Qwen3.5 122B A10B

Open-weight Qwen3.5 native vision-language MoE model (122B total, 10B active) with hybrid linear attention.

qwen3.5-122b-a10b
STABLEGet StartedView uptime
262,144 context
Released February 26, 2026
Starting at $0.40/M input tokens
Starting at $3.20/M output tokens
Streaming
Vision
Tools
Reasoning
Reasoning Budget
JSON Output
No ratings yetSign in to rate

Select Provider

All Providers for Qwen3.5 122B A10B

LLM Gateway routes requests to the best providers that are able to handle your prompt size and parameters.

Context: 262.1kQuant: bf16
Input
$0.4
/M tokens
Cache Read
/M tokens
Output
$3.2
/M tokens
Get Started
Context: 262.1k
Input
$0.4
/M tokens
Cache Read
/M tokens
Output
$3.2
/M tokens
Get Started

Frequently asked questions

What is Qwen3.5 122B A10B?

Open-weight Qwen3.5 native vision-language MoE model (122B total, 10B active) with hybrid linear attention. You can access it through LLM Gateway's OpenAI-compatible API with automatic provider routing, fallback, and cost analytics.

How much does Qwen3.5 122B A10B cost?

Pricing for Qwen3.5 122B A10B on LLM Gateway starts at $0.40 per million input tokens and $3.20 per million output tokens, depending on the provider. The pricing table above always reflects the current per-provider rates.

What is the context length of Qwen3.5 122B A10B?

Qwen3.5 122B A10B supports a context window of up to 262,144 tokens on its largest provider deployment.

Which providers serve Qwen3.5 122B A10B?

Qwen3.5 122B A10B is served by NovitaAI, Alibaba Cloud through LLM Gateway. Requests are automatically routed to the best available provider, with fallback when a provider has issues.

Does Qwen3.5 122B A10B support tool calling and structured outputs?

Qwen3.5 122B A10B supports tool (function) calling and soft JSON output, but NOT strict JSON output schema (its upstream provider does not enforce it).

When was Qwen3.5 122B A10B released?

Qwen3.5 122B A10B was released on February 26, 2026.