Q3 2026: Smarter Routing, Airside & Realtime
The Q3 2026 product roundup: smart and cache-aware routing, Airside, realtime voice and transcription, Lounge projects, coding tools, and stronger enterprise controls. Plus a public traffic snapshot: 50 million requests and 1.805 trillion tokens.

A coding session should not need the same model as a one-line classification. A voice app should not need a separate gateway. And getting an inference provider listed should not take a chain of sales calls.
In Q3, LLM Gateway shipped tools for each of those problems: smarter routing, realtime audio, a self-service provider console, and more ways to use the gateway from your editor, terminal, or browser. This Q3 2026 product roundup covers what shipped from July 1 through September 30.
Q3 2026 by the numbers
Across the platform, including gateway, DevPass, and Lounge traffic:
| Metric | July–September snapshot |
|---|---|
| Requests recorded | 50,036,094 |
| Total tokens processed | 1,805,087,667,040 — 1.805 trillion |
| Input tokens reported | 1,767,542,502,103 |
| Cached input tokens reported | 1,464,621,005,735 |
| Output tokens reported | 37,284,352,673 |
| Busiest recorded day | 1,576,727 requests on September 25 |
Measurement window: July 1–September 30, 2026, UTC. Snapshot retrieved September 30 at 16:35 UTC, across all usage modes and product types. September 30 is still in progress and aggregation can lag live traffic, so these are provisional quarter-to-date figures, not final quarter-end totals.
Request counts include recorded errors and gateway response-cache hits, not just successful upstream calls. Total tokens use the recorded total_tokens values; endpoint-specific accounting means input and output subtotals can differ from that total. Cached input is reported separately for context, not added again to total tokens.
Most-used models by requests
Model traffic is grouped by canonical model, combining its provider deployments. These are measured usage rankings, not a list of everything in the catalog.
| Rank | Model | Requests | Share of all requests |
|---|---|---|---|
| 1 | gemini-embedding-2 | 4,552,264 | 9.1% |
| 2 | deepseek-v4-flash | 3,785,467 | 7.6% |
| 3 | gpt-5.4-mini | 3,098,417 | 6.2% |
| 4 | gpt-5.6-luna | 3,068,951 | 6.1% |
| 5 | openai-moderation | 2,794,294 | 5.6% |
Gemini Embedding 2 led by request count. DeepSeek V4.1 Flash led by token volume, with 262,072,034,547 tokens. Embeddings and moderation appearing alongside chat models show why request count and token volume tell different stories.
Most-used providers by requests
| Rank | Provider | Requests | Share of all requests |
|---|---|---|---|
| 1 | openai | 11,565,098 | 23.1% |
| 2 | google-vertex | 6,384,358 | 12.8% |
| 3 | google-ai-studio | 5,503,630 | 11.0% |
| 4 | azure | 5,219,406 | 10.4% |
| 5 | deepseek | 3,399,602 | 6.8% |
OpenAI handled the most requests. DeepSeek led provider token volume, with 363,614,003,080 tokens. Shares use the full platform request total, not only these five providers; percentages are rounded to one decimal place.
Let the workload choose the model
Smart routing adds a new model: "smart" option. Choose the models it may use, then select either the cheapest eligible model or a classifier that rates the request's difficulty, task, and output type before choosing a model and reasoning effort.
- Your approved list stays the boundary. Organization defaults and project overrides decide the candidate set; the gateway does not quietly substitute a model outside it.
- Sessions can adapt between turns. The gateway reuses the current choice, then checks again when documented prompt caches expire or a periodic work scan is due. Harder work can move up; cost-driven changes account for rebuilding the cache. Tool-call continuations keep their current model.
- Decisions are visible. Request details show the classifier, candidates, selected model, effort, and why a choice changed.
- Existing
autocallers keep their behavior. Smart routing is a separate option, available to gateway organizations during beta, not DevPass yet.
Provider selection also became adaptive and cache-aware. Instead of ranking only uncached input prices, routing learns the project's recent cached-input and output mix and uses it to compare eligible providers. This improves the estimate for long prompts and coding sessions; it is not a promise that a particular prompt will hit a cache.
For teams that want explicit branching, dynamic routes add named flows invoked with dynamic/<name>: conditional rules, weighted splits, model fallbacks, classifier nodes, a visual or JSON editor, version history, and rollback. Dynamic routes are Enterprise-only.
Smart routing guide · Cache-aware routing · Dynamic routes
Put your inference service on the board with Airside

Airside is the new self-service console for model providers. Claim a carrier, register models, submit changes for review, and manage the listing without passing spreadsheets back and forth.
- Domain-verified claims connect a provider to its published endpoint or website.
- Fleet registration and preflight verification test declared capabilities against the upstream API before a model filing can be submitted.
- Multiple upstream formats let a provider choose the API its service actually implements.
- Crew roles and carrier branding make the console usable by a team, not just one account.
- Traffic reporting and listing controls give providers visibility into the requests their service receives.
A carrier's preflight test key is stored encrypted, reused for later verification runs, and can be replaced or removed in Settings. The resources hub adds listing guides and browser-based tools for preparing a service before submission.
Introducing Airside · Open Airside
More provider integrations and cloud options
Q3 expanded where requests can run without changing the gateway API:
- Inference providers: SCX.ai, Gonka24, Runware, Fireworks, RanoAI, and Consensus Protocol.
- Direct model APIs: new integrations with Meta and Atria.
- Cloud integrations: Baidu Qianfan International and Tencent Cloud added more deployment options.
- Claude on Microsoft Foundry: the
azure-anthropicintegration supports Anthropic Messages, streaming, provider cache-control passthrough, and server-side tool search. Foundry release notes. - Azure Priority processing: eligible Azure deployments now accept
service_tier: "priority"for chat completions and Responses. Request logs distinguish the requested tier from the tier Azure actually served. Priority is not included in coding plans. Azure Priority release notes.
Talk and transcribe through the same gateway
Realtime voice now uses an OpenAI-compatible WebSocket endpoint at wss://api.llmgateway.io/v1/realtime, bringing voice sessions under the same gateway authentication and usage controls as the rest of an application.
- Ephemeral client secrets let browser apps connect without exposing a long-lived gateway key.
- Speech-to-speech sessions keep the upstream realtime event protocol.
- Transcription-only sessions stream audio into text without requiring a model that speaks back.
- Lounge's realtime workspace includes Talk and Transcribe modes, live transcripts, Copy transcript, server voice-activity detection, and manual turn commits.
Realtime API docs · Try transcription in Lounge
Make coding tools easier to connect

DevPass Code launched as a terminal coding agent on npm. The wider CLI now removes much of the setup work for using other agents through the gateway:
llmgateway launchconfigures and starts supported coding agents, checking the key before launch.- Browser login works with an existing dashboard session, including enterprise SSO, instead of asking for a password in the terminal.
- Organization skills distribute shared instructions to supported agents; remote terminals can use a printed approval link and code.
- The official VS Code extension puts gateway models in Copilot Chat's model picker, with streaming and agent mode when the selected model supports tools. Both gateway and DevPass keys work.
- Agent attribution and analytics recognise GitHub Copilot and Empryo as first-class coding agents and add per-model usage breakdowns.
- Expanded integration guides cover more editor and terminal workflows.
DevPass gained Reset Passes for the weekly premium allowance, upgrade rollover and scheduled upgrades, optional overflow after the monthly allowance, a dedicated usage page, clearer reset timestamps, and saved-card removal after cancellation. Its new No AI training setting restricts routing to providers whose published policies meet that requirement, including fallback paths. It does not, by itself, disable gateway payload storage.
DevPass Code · CLI and organization skills · VS Code guide
Give chats a workspace in Lounge
The chat app became Lounge, with its own name, memberships, and workspace at lounge.llmgateway.io.
Projects bring a knowledge base, editable memory, shared instructions, and grouped chats into one place. Upload documents once, let them be indexed, and chat against them with source-file citations instead of re-uploading the same context every session. Project memory carries useful facts into new chats in the same workspace.
Lounge also brought realtime voice and transcription into the browser and continued improving its image, video, and audio workflows. Here is a generated-video example from September's Video Studio walkthrough:
AI-generated demonstration, not real footage. See the Video Studio walkthrough for the workflow.
Lounge projects and knowledge bases · Open Lounge
Use search, decisions, and video in your application
The gateway added API surfaces for more than text generation:
- Perplexity Search at
/v1/searchreturns ranked web results and extracted content, with domain, language, country, and recency filters. Use the results in your own agent or retrieval pipeline. - System One at
/v1/systemonereturns typed decisions: yes/no probabilities, choices with confidence, and scores. Applications can branch on values instead of parsing a prose answer. - AI SDK gateway-protocol compatibility lets existing SDK applications change the gateway base URL while keeping model-string resolution and supported web-search citations.
- Image quality controls add
xhighandmaxoptions through the images API, chat completions, and Lounge. - Reference-image generation adds multiple reference inputs on compatible ByteDance image deployments, mapping OpenAI-style image parts to the provider format.
- Version 4 of
@llmgateway/ai-sdk-providertargets AI SDK 7 and adds video generation, including job polling and content retrieval, alongside text, images, and tools. - TanStack AI integration guidance adds another documented path for using the gateway in an application.
Perplexity's Sonar integration also moved to its Agent API. The Pro Sonar models were retired during September; applications using them need to select a supported replacement from the live directory.
Search API · Typed decision API · AI SDK 7 video · Sonar migration
Popular Q3 model updates
- GPT-5.6 Luna
- GPT-6 Astra
- GPT-6 Sol
- GPT-6 Luna
- GPT-6.1 Sol
- Claude Fable 5.1
- Claude Opus 5.5
- Claude Sonnet 5.5
- DeepSeek V4.1 Flash
- Gemini 3.8 Flash
- Muse Spark 1.3
- Kimi K3
- GLM-5.3
- GLM-5.3 Flash
- Grok 4.7
- GPT Image 2.5 Sunburst
- GPT Image 2.5 Flare
- MiniMax H3 Max
Enforce privacy and team policy at the gateway
Q3 added stronger controls for both data handling and access:
- Zero data retention, Enterprise. Enforce compatible provider policies, Metadata Only gateway retention, disabled response caching, stripped provider prompt-cache markers, and
store: falseon the Responses API. Asynchronous video generation is unavailable under ZDR because its jobs require temporary output storage. - Provider headquarters restrictions, Enterprise. Route only through providers based in allowed countries; unknown headquarters fail closed.
- Compliance alerts, Enterprise. Watch blocked models and notify the selected recipients when a compliant route becomes available. Choose email or supported webhook channels.
- Project-scoped developers and organization teams, Enterprise. Apply project ceilings, per-developer limits, and shared IAM policies. Default teams and Microsoft Entra group sync over SCIM reduce manual onboarding.
- Per-member usage limits, all plans. Administrators can cap active keys and usage and set default developer limits without waiting for Enterprise team controls.
- Member analytics, Enterprise. Team administrators can inspect per-member requests, tokens, API keys, and model/provider breakdowns.
- Project-level guardrail overrides, Enterprise. Apply the appropriate policy to each workload, with better redaction behavior.
- Enterprise license expiry warnings. Licensed Enterprise dashboards now warn before license expiry.
- Hash-only gateway key storage. New and rotated secrets are shown once; stored HMAC fingerprints replace recoverable gateway API-key secrets. Provider credentials remain encrypted for upstream authentication.
Compliance controls · Team and access controls
Understand usage without leaving your workflow
- Dashboard comparison periods overlay a previous or custom range on the usage chart instead of requiring two exports.
- Key usage and reset times in the Master Keys API and the embeddable SDK's
getBalance()report consumed usage alongside configured limits and the current window's reset time. - API-key capacity indicators flag keys approaching or reaching their limits, with filters to find them.
- Timezone-aware analytics align daily buckets with the browser's local day; timestamp displays can switch between local time and UTC.
- MCP usage analytics add read-only usage summaries and activity search, scoped to the connected project. Owners and admins see that project's usage; developers see only their own keys.
- The organization Models directory combines catalog and custom models with compliance eligibility when an Enterprise policy is active.
- Clearer key and model details add provider-key descriptions, creator and project names in Master Keys, model lifecycle badges, and sortable provider tables.
MCP analytics guide · Usage comparisons
Keep integrations running
The less visible work matters when a session runs for hours:
- DeepSeek predecessor migration. DeepSeek V4.1 Flash launched on September 10. Its first-party V4 Flash and V4 Flash Vision predecessor mappings retired that day. Update integrations pinned to those retired mappings to
deepseek/deepseek-v4.1-flash; third-party V4 Flash routes were not affected by that migration. - SSE keepalives and terminal error events reduce idle disconnects and make interrupted Anthropic and Responses streams explicit.
- Client-managed prompt caching separates automatic gateway markers, preserved client markers, and fully disabled caching.
- Trust tiers let limits grow with account age or qualifying usage. Settings → Limits shows the current tier, per-endpoint limits, usage caps, and what is needed for the next tier. Trust tier release notes.
- Concurrent request limits and overload protection bound in-flight work across inference endpoints, alongside clearer account-limit indicators and notifications.
- Safer content-filter handling distinguishes provider safety blocks from other failures and checks the relevant turn or streaming chunk.
- More provider and regional coverage, including additional cloud deployments and service-tier support, arrived alongside ongoing model lifecycle updates.
The catalog also improved how it presents non-token-priced models, with per-image, per-second, and per-character filters. For current model availability, capabilities, regions, and provider coverage, use the live models and providers directories rather than a quarter-end list that goes stale.
Smaller improvements across the products
Q3 also shipped downloadable invoices and credit notes, crypto top-ups, self-service refund flows, branded end-user receipts in the embeddable SDK, clearer purchase and subscription controls, and more setup and migration guides.
The support widget gained suggested starter questions, conversation ratings, and one-click human escalation with email follow-up.
For dated release notes and the individual updates behind this post, browse the changelog. For the previous quarter, see the Q2 2026 roundup.
- Try LLM Gateway for your next application.
- Read the API docs to connect an existing integration.
- Explore smart routing for request-by-request model selection.
Frequently asked questions
- What shipped on LLM Gateway in Q3 2026?
- Smart model selection, adaptive cache-aware provider routing, Airside's self-service provider console, realtime voice and transcription, Lounge projects, a VS Code extension, CLI browser login and organization skills, plus expanded enterprise compliance and team controls.
- How much traffic did LLM Gateway process in Q3 2026?
- The September 30 snapshot records 50,036,094 requests and 1,805,087,667,040 total tokens for July through September. September 30 is still in progress, so these are provisional quarter-to-date figures, not final quarter-end totals.
- Which model and provider were used most in Q3 2026?
- By request count, Gemini Embedding 2 led models and OpenAI led providers. By total tokens, DeepSeek V4.1 Flash led models and DeepSeek led providers. Rankings use the same provisional July–September snapshot.