Back to blog

New AI Models This Week: Haiku 5.5, Mistral Large 4, GLM 5.3 Fast

Every new AI model released between September 28 and October 11, 2026 and already live on LLM Gateway: Claude Haiku 5.5 and Sonnet 5.5, GPT-6.1 Sol, Mistral Large 4, GLM 5.3 Fast, Reka Edge 2603, Gemini Nano Banana 2.1, Hy Image 3.5 Preview, Ling 3.1 Flash and Grok Imagine Video 1.5 Lite, with prices, context windows and the model ID to call.

By LLM Gateway

A glowing calendar page with a sparkle on the central chip of a dark circuit board, surrounded by rocket, film clapperboard, paintbrush, chat bubble, brain, price tag, stopwatch and star icons, representing the new AI models released this week

Ten new AI models landed in the LLM Gateway catalog in the last two weeks, from a $0.10-per-million-token Claude to a 1.05-trillion-parameter open-weight Mistral. Each text model is callable today with the same chat request you already send; only the model string changes. Image models go through /v1/images/generations and the video model through /v1/videos. Prices are per million tokens unless noted, as listed on each model page on October 11, 2026.

New AI models at a glance

Released Model Family Input / output price Context Notes
Oct 9 Reka Edge 2603 Reka $0.10 / $0.10 16K Small vision model with tool calling, via Reka AI
Oct 8 Grok Imagine Video 1.5 Lite xAI $0.02/s (480p), $0.03/s (720p) — Text- or image-to-video
Oct 7 Claude Haiku 5.5 Anthropic $0.10 / $0.50 up to 100K-token prompts; $0.50 / $2.50 above 1M Tools, vision, adaptive thinking
Oct 7 GLM 5.3 Fast Z.ai $2.80 / $8.80, cached input $0.56 1M Fast GLM 5.3 variant, 128K output, served by SCX.ai
Oct 6 Mistral Large 4 Mistral $1.36 / $4.18 1M Open-weight MoE, 49B active / 1.05T total, vision
Oct 6 Gemini Nano Banana 2.1 Google $1.50 / $7.50, $30 per 1M image output tokens 131K Image generation and editing, 1K–4K output
Oct 5 Hy Image 3.5 Preview Tencent $1.60 per 1M image output tokens 100K Image generation and editing, up to 20 reference images
Oct 2 Ling 3.1 Flash InclusionAI Free (Novita); $0.06 / $0.18 on DeepInfra 262K Hybrid reasoning MoE, 560B total / 25B active, tools
Sep 29 GPT-6.1 Sol OpenAI $2.00 / $10.00 1.05M Coding, computer use, reasoning
Sep 28 Claude Sonnet 5.5 Anthropic $2.00 / $10.00 1M Adaptive thinking, vision, tools

Claude Haiku 5.5: a 1M-context model at $0.10 per million input tokens

Anthropic positions Haiku 5.5 as its fastest model for high-volume, latency-sensitive work: classification, routing, extraction and subagent tasks. It keeps the adaptive thinking and vision capabilities of Sonnet 5.5 at a twentieth of the price for prompts up to 100K tokens, with cached input at $0.01 per million tokens and a 128K output limit. Above 100K prompt tokens the rate steps up to $0.50 input / $2.50 output ($0.05 cached), so the long-context tier is priced like a quarter of Sonnet rather than a twentieth. On LLM Gateway it is served by Anthropic and AWS Bedrock, so a bare claude-haiku-5-5 request fails over between them.

1curl https://api.llmgateway.io/v1/chat/completions \2  -H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \3  -H "Content-Type: application/json" \4  -d '{5    "model": "claude-haiku-5-5",6    "messages": [{"role": "user", "content": "Label this support ticket: billing, bug, or feature request."}]7  }'

If you run a router-and-worker pattern, Haiku 5.5 as the router and Sonnet 5.5 or GPT-6.1 Sol as the worker is the obvious pairing this week.

Mistral Large 4: an open-weight flagship with a 1M context

Mistral Large 4 is a granular mixture-of-experts model with 49B active and 1.05T total parameters plus a 1.6B vision encoder, and the weights are open. At $1.36 input / $4.18 output it sits between Haiku and Sonnet on price while matching their 1M-token context. Cached input is $0.14. It supports tool calling, so it drops into agent loops as a cheaper alternative to the closed flagships. Model ID: mistral-large-4.

Gemini Nano Banana 2.1 and Hy Image 3.5: two image models in one week

Google's Nano Banana 2.1 succeeds Nano Banana 2 with sharper 1K to 4K output, better text rendering, multi-turn character consistency and search grounding. Text in and out is priced like a Flash-tier model ($1.50 / $7.50) and image output at $30 per million image tokens.

Tencent's Hy Image 3.5 Preview is a unified generation-and-editing model on the Hy Image 3.0 MoE base. It accepts up to 20 reference images, outputs up to 4K and renders Chinese and English text well. Image output is $1.60 per million image tokens.

Both are available through /v1/images/generations and the Lounge image studio.

GPT-6.1 Sol and Claude Sonnet 5.5: the new mid-tier pair

The two releases from the week before this one are priced identically, $2.00 input / $10.00 output, with cached input at $0.10, and both offer a 1M-class context (1.05M for Sol) and 128K output. OpenAI describes Sol as near-Astra performance at lower cost for complex coding, computer use and professional work; Anthropic describes Sonnet 5.5 as its best combination of speed and intelligence. If you are choosing between them for a coding agent, run both through the same prompts on the compare view before committing; usage share across LLM Gateway customers is on the rankings page.

GLM 5.3 Fast and Reka Edge 2603: two Airside listings

Not every new model comes from a hyperscaler. GLM 5.3 Fast is a faster variant of Z.ai's GLM 5.3 with a 1M context and 128K output, listed on LLM Gateway by the carrier SCX.ai at $2.80 input / $8.80 output with cached input at $0.56; it is text-only and does not expose tool calling on this listing. Reka Edge 2603 is Reka's small vision model, 16K context, with tool calling, at $0.10 per million tokens in and out from Reka AI. Both arrive through Airside, the carrier program that lets any provider list models on the gateway, which is why they appear on the timeline before the open-source catalog picks them up. Model IDs: glm-5.3-fast and reka-edge-2603.

Free this week: Ling 3.1 Flash

InclusionAI's Ling 3.1 Flash is a hybrid-reasoning MoE (560B total, 25B active) with tool calling and a 262K context, and it is free on LLM Gateway via Novita (DeepInfra also serves it at $0.06 / $0.18 if you want a paid fallback). It is a good default for evals and CI jobs where you want a capable model at no cost. Model ID: ling-3.1-flash.

Video: Grok Imagine Video 1.5 Lite

xAI's lightweight video model generates clips from a text prompt or an input image at $0.02 per second (480p) or $0.03 per second (720p), plus $0.01 per input image, from us-east-1 and us-west-2. Call it through /v1/videos.

Keep up with new AI models

Frequently asked questions

Which new AI models were released this week?
Between October 2 and October 9, 2026: Ling 3.1 Flash (InclusionAI, free), Hy Image 3.5 Preview (Tencent), Mistral Large 4, Gemini Nano Banana 2.1 (Google), Claude Haiku 5.5 (Anthropic), GLM 5.3 Fast (Z.ai via SCX), Grok Imagine Video 1.5 Lite (xAI) and Reka Edge 2603 (Reka). The week before brought Claude Sonnet 5.5 and GPT-6.1 Sol.
How much does Claude Haiku 5.5 cost?
$0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens, with cached input at $0.01 per million. Prompts over 100K tokens are billed at $0.50 input / $2.50 output and $0.05 cached input. It has a 1M-token context window and supports tools, vision and adaptive thinking.
How do I use a new model on LLM Gateway?
Set the model field of your request to the ID in the table, for example claude-haiku-5-5 or mistral-large-4, and send it to https://api.llmgateway.io/v1/chat/completions with your LLM Gateway API key. Image models are called through /v1/images/generations and video models through /v1/videos. New models are added to the catalog as providers publish them, usually on release day.