GPT Image 2.5 Is Here
Generate and edit images with OpenAI's GPT Image 2.5 Sunburst and Flare. Both add xhigh and max quality settings through the images API, chat completions, and Playground.
Read more about GPT Image 2.5 Is Here →Stay up to date with the latest features, improvements, and fixes in LLM Gateway.
Generate and edit images with OpenAI's GPT Image 2.5 Sunburst and Flare. Both add xhigh and max quality settings through the images API, chat completions, and Playground.
Read more about GPT Image 2.5 Is Here →The /v1/realtime WebSocket now opens transcription-only sessions: live speech-to-text with no speech model in the loop, billed per minute or per token against the transcription model alone. Lounge gets a Transcribe mode on the same page.
Read more about Realtime Transcription Sessions →@llmgateway/ai-sdk-provider 4.0 targets Vercel AI SDK 7 and adds llmgateway.video() for experimental_generateVideo: the provider submits a gateway video job, the SDK polls it, and you get the MP4 bytes back, with image-to-video, first and last frames, reference inputs, and audio control. Stay on 3.x for AI SDK 6.
Read more about AI SDK 7 Provider with Video Generation →Airside now runs a preflight against your endpoint before a model can be filed: every capability you declare is probed live, and the results become the capability badges developers see. Carriers can also pick the upstream API format per model, invite crew by email, file per-model fares, and rename their carrier under review.
Read more about Airside: Model Verification and Crew Invites →Smart routing now ranks providers on what a cached workload actually pays. For prompts of 5,000 tokens or more, the price score blends each provider's cached input rate at an assumed 70% hit rate and weights output at a 20% output-to-input ratio, and context-length pricing tiers are resolved from the prompt size before scoring. Both assumptions are tunable per project on the Enterprise plan.
Read more about Cache-Aware Provider Selection →Sign in to the LLM Gateway CLI from your browser or enterprise SSO instead of typing a password, and pull your organization's shared skills into Claude Code, Codex, OpenCode, Cursor, and other agents with one command. Browser login works on every plan; organization skills are available on the Enterprise plan.
Read more about CLI Browser Login and Organization Skills →The dashboard usage chart now overlays a comparison period, whether the previous period, a chosen week or month, or an exact custom range, and switches between a total view and a token-cost breakdown. Bars are grouped and stacked with an explicit legend, and rankings gain a hover-isolating legend and an all-time top apps panel.
Read more about Compare Usage Across Periods →Gemini 3.8 Flash lands on Google AI Studio and Vertex AI with a 1M context at $0.75/M input, Meta's Muse Spark 1.3 arrives with a 1M context and a $0.10/M Contributor tier, and Kimi K3, GLM-5.3 Flash, and Qwen3.8 Flash pick up new deployments across Runware, Novita, SCX.ai, and Alibaba Cloud.
Read more about Gemini 3.8 Flash, Muse Spark 1.3 & More Models →The LLM Gateway MCP server gains get-account, get-usage, and get-usage-breakdown, so Claude Code, Codex, Cursor, or any MCP client can check spending limits, request and token totals, costs, trends, and your most-used providers, models, coding apps, and API keys without opening the dashboard.
Read more about Usage Analytics in the MCP Server →The models directory filters and sorts image, video, and speech models by their real unit prices and shows retirement status as chips; model pages sort providers by price, speed, or context; Seedream 5.0 Pro accepts up to 10 reference images; DevPass shows exact renewal and reset times; and Enterprise licenses warn 90 days before expiry.
Read more about Per-Unit Price Filters, Seedream References & More →Enterprise organizations can now enforce zero data retention across provider routing, LLM Gateway storage, response caches, and the Responses API. Conflicting retention and caching settings are blocked before they can weaken the policy.
Read more about Zero Data Retention Controls →Airside is the new carrier console where LLM providers list themselves on the gateway: claim your provider by verifying your company domain, register models, file prices for review, and tune the margin and discounts that win routed traffic. Listing costs a one-time $2,500 fee per provider company and goes live once we approve your claim.
Read more about Airside: Self-Serve Provider Listings →Anthropic's Claude Fable 5.1 is live on the gateway with a 1M context window, always-on adaptive thinking, stronger long-horizon agentic work, and cache reads at a quarter of Fable 5's price.
Read more about Claude Fable 5.1 →Claude now routes through Microsoft Foundry under a new azure-anthropic provider. The models directory gained lifecycle status filters, model pages sort providers by price, speed, or context, and API key lists flag the keys that are near or at their limit.
Read more about Claude on Foundry, Model Status & More →DevPass subscribers can now remove their saved card after canceling, or whenever a subscription is inactive. The card details leave Stripe while a privacy-safe fingerprint remains to enforce the one-card-per-account rule.
Read more about Remove Your DevPass Payment Card →DevPass subscribers can now restrict routing to providers that explicitly state API inputs are not used for training. Unknown policies fail closed, and retries or fallbacks never escape the setting. Available on every active DevPass tier.
Read more about No AI Training for DevPass →Group developers under one shared policy — a project ceiling, per-developer budgets, and IAM rules — instead of configuring each person by hand. Microsoft Entra groups map onto teams over SCIM, and a default team catches everyone who joins without one. Available on the Enterprise plan.
Read more about Organization Teams and Directory Sync →Gateway API keys are now stored only as keyed HMAC-SHA-256 fingerprints and provider credentials only as AES-256-GCM ciphertext. Every plaintext read path is gone, so a key's secret is visible exactly once per issuance — at creation or roll.
Read more about Hash-Only API Key Storage →Every organization now has a fleet-wide budget of concurrent in-flight requests that scales with your trust tier — bounding the long-running streams a per-minute limit can't see. Over-budget requests get a retryable 429, and a momentarily saturated gateway sheds load with a 529 instead of queueing until it degrades.
Read more about Concurrent Request Limits and Overload Protection →Provider keys now take a description that follows them through key lists and request drilldowns, the master API returns each key's creator and project by name, and the catalogue picks up Grok 4.6 on AWS Bedrock and Vertex AI plus GLM-5.2 Turbo and an Australian region on SCX.ai.
Read more about Provider Key Descriptions, Grok 4.6 & More →Provider cache writes is now a three-way setting. The new Client-managed mode forwards the cache markers your client sends and never adds any of its own, so one API key can serve a coding agent that manages its own caching alongside traffic that should not pay the cache-write premium.
Read more about Client-Managed Prompt Caching →Every organization now has transparent per-endpoint rate limits, daily/monthly spend caps, and top-up allowances that scale automatically with account age or lifetime usage — all visible on the new Settings → Limits page. Enterprise organizations have no limits at all.
Read more about Trust Tiers: Limits That Grow With Your Account →LLM Gateway now speaks the AI SDK's own gateway protocol, so an app built on the Vercel AI Gateway runs here with one line changed. Bare model strings, provider-native web search with citations, and the model picker all keep working.
Read more about Drop-In For The Vercel AI Gateway →You can now top up your credits balance with crypto. Open the top-up dialog, take the checkout option below the amount, and pick crypto on the Stripe Checkout page — the credits land on your account exactly like a card payment.
Read more about Pay For Credits With Crypto →Define named, versioned routing flows — conditions, A/B splits, and model targets with provider fallback — and invoke them by putting dynamic/<name> in the model field. Build them visually or in JSON, publish immutable versions, and roll back instantly. Available on the Enterprise plan.
Read more about Dynamic Routes →DevPass no longer has to stop at 100%: opt into pay-as-you-go overflow and, once your monthly allowance is used, requests keep flowing from a credits balance billed at provider rates. Top up from the dashboard with your saved card, set auto-reload so the balance refills itself, and track it all on the new Usage page.
Read more about DevPass Pay-As-You-Go Overflow →Starting Friday, August 7, 2026, the /v1/moderations endpoint is billed at a flat $0.00001 per successful request — no token metering, no per-model rates, and no charge for failed or retried attempts. Because it becomes paid, it also starts requiring a credit balance: top up before August 7 or moderation stops working for organizations with no credits.
Read more about Moderations Pricing From August 7 →The dashboard's Custom Models page is now a full Models directory: every catalog model plus your organization's custom models in one searchable table, with per-model compliance eligibility so your team sees exactly what they can route to.
Read more about Org Models Directory →OpenAI cut GPT-5.6 Terra to $2.00/$12.00 and Luna to $0.20/$1.20 per 1M tokens — Terra is 20% cheaper, Luna 80%. The new rates are live on LLM Gateway and apply automatically to every request, including cached input, cache writes, and long-context pricing.
Read more about GPT-5.6 Terra & Luna Price Cuts →We partnered with Runware.ai to serve six open-source flagships — gpt-oss-120b, Gemma 4, DeepSeek V4 Pro and Flash, Kimi K2.6, and GLM 5.2 — at 30% off. The discount applies automatically to everything routed through Runware until September 9.
Read more about Runware Launch: 30% Off Open-Source Models →Speech-to-speech conversations now run through LLM Gateway over the OpenAI-compatible /v1/realtime WebSocket endpoint — with ephemeral client secrets for browsers, per-token audio billing in your activity feed, and a voice call surface in Lounge.
Read more about Realtime Voice API →Our consumer chat app has a name: Lounge by LLM Gateway — the members' lounge for AI. Same app, same prices, same chat.llmgateway.io, with a full visual identity: membership pricing, boarding-pass plan cards, and a proper wordmark.
Read more about Chat Is Now Lounge →The LLM Gateway CLI can now start any supported coding agent pre-wired to the gateway — Claude Code, OpenCode, Codex CLI, DevPass Code, and eight more. One command, one API key, 200+ models, and every request tracked in your dashboard.
Read more about Launch Any Coding Agent from the CLI →DevPass upgrades now roll your unspent allowance into the new tier — or schedule the switch for your next renewal. Plus two new providers including SCX.ai's Turbo inference (up to 4x faster), Gemini 3.6 Flash and 3.5 Flash Lite, Gemini TTS on two providers, and Empryo as a first-class coding agent.
Read more about Upgrade Rollover, New Providers & Gemini TTS →Hit your weekly premium allowance mid-sprint? A Reset Pass instantly restores the full allowance and starts a fresh 7-day window. Pro includes 1 pass per cycle and Max includes 2; extra passes are a one-time purchase from the DevPass dashboard.
Read more about DevPass Reset Passes →Restrict routing to providers headquartered in the countries you approve. Pick allowed countries on the compliance page and the gateway blocks any provider based elsewhere — before a request leaves the gateway. Available on Enterprise.
Read more about Provider Headquarters Compliance Filter →OpenAI's new GPT-5.6 family is live on LLM Gateway from day one: Sol for frontier reasoning at $5/$30 per 1M tokens, Terra balancing intelligence and cost at $2.50/$15, and Luna for high-volume workloads at $1/$6 — all with a 1.05M-token context window, reasoning, vision, tool calling, and web search.
Read more about GPT-5.6 Sol, Terra & Luna Now Available →Cap any teammate's spend and API keys with per-member budgets, give contractors project-scoped developer access on Enterprise, download a PDF invoice for any purchase, and see analytics bucketed in your own timezone. Plus Claude Sonnet 5 at introductory pricing, Claude Fable 5 back online, and DevPass Code on npm.
Read more about Per-Member Budgets, Developer Role & PDF Invoices →Roll cost, requests, and tokens up across every project in your organization, then break the spend down by model, project, or API key over any date range. Read from pre-aggregated rollups, so it stays fast on any window. Available to owners and admins on the Enterprise plan.
Read more about Organization-Wide Analytics →See exactly where your spend goes. Every project gets a Cost by Model analytics page, every API key gets its own statistics page, and Enterprise orgs get per-member usage breakdowns — all on the date-range picker you already use. Member analytics are available on Enterprise.
Read more about Usage Analytics by Model, Key and Member →Bring any OpenAI-compatible model under full cost tracking. Define pricing, context limits, and capabilities per custom provider key — so requests through custom providers get billed, enforced, and reported just like a built-in model. Available on Enterprise.
Read more about Custom Model Catalog for Custom Providers →Restrict routing to providers that meet your compliance requirements — SOC 2, ISO 27001, GDPR, no prompt training, no prompt logging. Requests to non-compliant providers are blocked before any data leaves the gateway. Available on Enterprise.
Read more about Provider Compliance Policies →Steer multi-provider routing with a new routing field — auto, price, throughput, or latency — per request or as a per-project default. Each strategy still falls back when the top pick has bad uptime.
Read more about Routing Strategies: Cheapest, Fastest & Defaults →We've temporarily suspended access to Claude Fable 5 across all providers while we work through a usage-policy matter with Anthropic. Requests to the model now return a clear error, and routing automatically skips it.
Read more about Claude Fable 5 Access Suspended →Monthly Chat subscriptions from $9, Flex and Priority service tiers, sandbox test keys for the LLM SDK, a no_training model filter, public DevPass profiles, and a stack of product polish.
Read more about Chat Plans, Service Tiers, SDK Sandbox & More →Anthropic's next-generation Claude Fable 5 lands with 1M context, Reve joins as a new image provider, xAI's Grok Imagine Video 1.5 turns images into 15-second clips, and NVIDIA's Nemotron 3 Ultra 550B arrives.
Read more about Claude Fable 5, Reve Image Gen & More New Models →Text-to-speech is live: nine models from ElevenLabs, OpenAI, and Gemini behind the OpenAI-compatible /v1/audio/speech endpoint, plus a new Audio Studio in the Playground to compare voices side by side.
Read more about Speech Generation + Audio Studio →Give any API key a time-to-live when you create it — minutes, hours, or days. Expired keys are disabled automatically, and you can bring them back online anytime with a fresh expiration.
Read more about API Key Expiration (TTL) →Claude Opus 4.8 lands with a 1M context window, Sonnet 4.6 gets 1M context too, and Qwen3.7 Max, Grok Build 0.1, Kimi K2.6, and GLM-5.1 join the gateway.
Read more about Claude Opus 4.8 + a Wave of New Models →Share a referral link and reward new signups, toggle provider cache writes per project, see cache and audio costs in your logs, plus a faster Chat and more.
Read more about Referrals, Cache Controls & Product Updates →Pin conversations to one provider for warm caches with x-session-id, route Bedrock by region, and reach Gemini via Google AI Studio plus Vertex embeddings.
Read more about Smarter Routing: Sticky Sessions, Bedrock Regions & More →Send PDFs and text-family documents to Gemini models via the OpenAI-compatible `file` content block.
Read more about Document Reading (PDFs & more) →ByteDance Seedance video models land in the gateway, Chat gets pinning and cross-org sharing, plus vertex-anthropic, grok-4.20, and a stack of fixes.
Read more about Seedance Video Models, Pinned Chats, Sharing Across Orgs & More →Turn text into vectors for semantic search, clustering, and RAG — through the same gateway you already use for chat.
Read more about OpenAI-Compatible Embeddings →DevPass by LLM Gateway is live on Product Hunt today. One subscription, every coding model, three flat prices.
Read more about DevPass is live on Product Hunt →Sessions are now Agents — monitor your AI coding agents, track costs per agent, and drill into individual sessions.
Read more about Sessions Rebranded to Agents →Route requests to regional providers, protect your apps with built-in content moderation, enforce API key rate limits, and explore new models.
Read more about Multi-Region Routing, Content Filters & More →Generate videos via the API, track conversations with sessions, and more — plus new models and providers.
Read more about Video Generation, Sessions & More →Access OpenAI's most capable models — GPT-5.4 for complex professional work and GPT-5.4 Pro for smarter, more precise responses — with 1.05M context windows and reasoning support.
Read more about GPT-5.4 and GPT-5.4 Pro Now Available →A dedicated Image Studio in the Playground for gallery-based generation with multi-model comparison, an OpenAI-compatible /v1/images/edits endpoint, and a wave of image generation improvements.
Read more about Image Studio, Image Edits API & More →When a provider fails, LLMGateway now automatically retries your request on another provider. Every attempt is logged with full routing visibility, so you always know what happened.
Read more about Automatic Retry & Fallback with Full Routing Transparency →Build AI-powered applications faster with pre-built agents, production-ready templates, and a new CLI tool for scaffolding projects.
Read more about AI Agent skills, Agents, Templates & CLI →New unified reasoning object for precise control over reasoning models. Specify exact token budgets with max_tokens or use effort levels — all in one consistent API.
Read more about Unified Reasoning Configuration →Ship faster with Dev Plans — AI-powered development planning now in beta. Plus native web search for real-time data, MiniMax provider, structured outputs for Anthropic & Perplexity, and a redesigned models experience.
Read more about Dev Plans, Native Web Search, and MiniMax Provider →Track all organization activity with comprehensive audit logs. See who did what, when, and to which resource — available for Enterprise customers.
Read more about Enterprise Audit Logs →Protect your LLM usage with content guardrails. Detect and block prompt injections, PII, secrets, and more — available for Enterprise customers.
Read more about Enterprise Guardrails →We're simplifying our pricing. All paid subscription features are now free for everyone — BYOK, team management, 30-day data retention, and more.
Read more about Pro Features Now Free for Everyone →Introducing Alibaba Cloud's Qwen Image model family - powerful models for text-to-image generation and image editing, now available in four variants: Qwen Image, Qwen Image Max, Qwen Image Max 2025-12-30, and Qwen Image Plus.
Read more about Alibaba Cloud Qwen Image Models: Advanced Image Generation and Editing →New Cerebras provider with six high-performance models, including GPT-OSS 120B and Qwen 3, now available through LLM Gateway.
Read more about Cerebras: Ultra-Fast Inference with 6 New Models →Google's latest Gemini 3 Pro Preview is now available with an exclusive 20% launch discount, featuring 1M context window and prompt caching.
Read more about Gemini 3 Pro Preview: 20% Off Launch Discount →Introducing Sherlock Dash Alpha and Sherlock Think Alpha (Grok 4.1) - free stealth models with 1.8M context, reasoning, vision, and advanced capabilities.
Read more about Sherlock: Two New Stealth Alpha Models →CanopyWave brings Kimi K2 Thinking to LLM Gateway with an exclusive 75% discount.
Read more about CanopyWave: 75% Off Kimi K2 Thinking →CanopyWave brings Qwen3 Coder, MiniMax M2, and GLM-4.6 to LLM Gateway with an exclusive 75% discount on all three models.
Read more about CanopyWave: 3 New Models with 75% Off →Added support for Moonshot AI's Kimi K2 Thinking model with 262K context window, advanced reasoning capabilities, and prompt caching for cost-effective thinking tasks.
Read more about Kimi K2 Thinking Model Support →Save on all Z.ai models with 10% off and get 20% off all Google models through LLM Gateway.
Read more about Z.ai 10% Off & Google 20% Off All Models →Added native support for AWS Bedrock, Google Vertex AI, and Microsoft Azure.
Read more about AWS Bedrock, Google Vertex AI and Microsoft Azure →Invite teammates, assign roles (Owner, Admin, Developer), track included seats, and add more seats as your org grows. Pro includes team management; Enterprise adds SSO/SAML, SCIM, audit logs, and advanced permissions.
Read more about Team Members: Roles, Seats, and Access Controls →Exclusive partnership with CanopyWave brings massive 90% discount on DeepSeek v3.1, making advanced reasoning capabilities more accessible than ever.
Read more about CanopyWave Partnership: 90% Off DeepSeek v3.1 →Added support for Anthropic's Claude Sonnet 4.5
Read more about Claude Sonnet 4.5 Model Support →Added support for Grok 4 Fast Reasoning, and Grok 4 Fast Non-Reasoning models via xAI provider.
Read more about Grok 4 Fast Models: Flagship and Fast Variants Now Available →Added support for qwen-max, qwen-max-latest, qwen-plus-latest, qwen-flash, qwen-vl-max, qwen-vl-plus and the new Qwen3 Next 80B A3B Instruct and Thinking models.
Read more about New Alibaba Qwen Models: Qwen3 Next, Max, Plus, Flash, Vision →Configure Claude Code to use any LLM model through LLMGateway's unified API with simple environment variable setup.
Read more about Claude Code Configuration Now Supported →Access Alibaba's powerful Qwen3 Max model with 256K context window, advanced reasoning capabilities, vision support, and function calling - all at competitive pricing.
Read more about Qwen3 Max Model Now Available →Generate stunning images with Google's Gemini 2.5 Flash Image Preview - our first image generation model with 32.8k context window and competitive pricing.
Read more about Introducing Our First Image Generation Model: Gemini 2.5 Flash Image Preview →Added support for DeepSeek's latest v3.1 model with 128K context window and competitive pricing for advanced reasoning capabilities.
Read more about DeepSeek v3.1 Model Support →Browse and compare 100+ AI models with advanced filtering, plus access Meta Llama 3.1 70B Instruct FP8 completely free through our CloudRift partnership.
Read more about New Models Directory & Free Llama 3.1 70B via CloudRift →Set individual credit limits for API keys to better control spending and prevent unexpected overages.
Read more about API Key Usage Limits & Credit Controls →Get instant access to OpenAI's powerful new GPT-5 model family including gpt-5, gpt-5-mini, gpt-5-nano, and gpt-5-chat-latest with 400k context windows.
Read more about GPT-5 Model Family Now Available →Released v2.0 of our @llmgateway/ai-sdk-provider npm package with improved Vercel AI SDK integration and simplified model access.
Read more about AI SDK Provider v2.0 Released →Added support for Claude 4.1 models including claude-opus-4-1, claude-opus-4-20250514, and claude-sonnet-4-20250514 via Anthropic provider.
Read more about Claude 4.1 Models: Opus and Sonnet Now Available →Added support for GPT-OSS-120B and GPT-OSS-20B models via Groq, offering powerful open-source alternatives with extensive context windows and competitive pricing.
Read more about New GPT-OSS Models: 120B and 20B via Groq →We’ve moved from TanStack Start to Next.js. Here’s why it matters
Read more about Next.js migration →Added support for Cloudrift, Moonshot AI and Novita AI providers, both offering the powerful kimi-k2 model with extensive context windows and competitive pricing.
Read more about New Providers: Cloudrift, Moonshot AI and Novita AI Support →Added support for Groq and xAI providers with their latest models, plus credits are now always visible in the sidebar for easy access.
Read more about New Providers: Groq and xAI Support + Always-Visible Credits →Introducing organizations and projects for clearer controls and statistics.
Read more about Dashboard UI Improvements & Project Context →Bring your own LLM provider keys or use credits with reduced gateway fees (2.5% vs 5%). Includes premium analytics, higher rate limits, and priority email support.
Read more about Pro Subscription Launch →New and improved self-hosting documentation for teams and enterprises looking to deploy LLM Gateway on their own infrastructure.
Read more about Self-Hosting Just Got Easier →Massive savings with Deepseek models and the arrival of Mistral models for all users. Discover new performance benchmarks at lower costs.
Read more about Deepseek Discount + Mistral Joins the Lineup →The unified AI gateway is here! Access 30+ models from 8 providers through one OpenAI-compatible API with transparent pricing and powerful analytics.
Read more about LLM Gateway v1.0 Launch →