Support

AI-powered help

Welcome!

Please introduce yourself before we start.

    Long Context Models

    Models with context windows of 200K tokens or more — up to 2M — for whole-codebase and multi-document workloads

    Compare
    Use Case
    Capabilities
    Provider
    Status
    Input Price ($/M tokens)
    Output Price ($/M tokens)
    Context Size (tokens)
    136/258
    Models
    41/55
    Providers
    97
    Vision Models (filtered)
    126
    Tool-enabled (filtered)
    0
    Free Models (filtered)
    Features
    AWS Bedrock(global)
    claude-fable-5
    $10.00$50.00$1.00
    AWS Bedrock
    claude-fable-5
    $10.00$50.00$1.00
    Nebius AI
    nemotron-3-ultra-550b
    $1.00$3.00—
    DeepInfra
    nemotron-3-ultra-550b
    $0.50$2.20$0.10
    Alibaba Cloud(cn-beijing)
    qwen3.7-plus
    $0.40$1.60$0.08
    Alibaba Cloud(singapore)
    qwen3.7-plus
    $0.40$1.60$0.08
    Alibaba Cloud(us-virginia)
    qwen3.7-plus
    $0.40$1.60$0.08
    Alibaba Cloud
    qwen3.7-plus
    $0.40$1.60$0.08
    Alibaba Cloud(eu-frankfurt)
    qwen3.7-plus
    $0.28$1.10$0.06
    Nebius AI
    minimax-m3
    $0.30$1.20—
    Together AI
    minimax-m3
    $0.30$1.20$0.06
    MiniMax
    minimax-m3
    $0.60$2.40$0.12
    AWS Bedrock(au)
    claude-opus-4-8
    $5.50$27.50$0.55
    Anthropic
    claude-opus-4-8
    $5.00$25.00$0.50
    AWS Bedrock
    claude-opus-4-8
    $5.00$25.00$0.50
    AWS Bedrock(global)
    claude-opus-4-8
    $5.00$25.00$0.50
    AWS Bedrock(jp)
    claude-opus-4-8
    $5.50$27.50$0.55
    AWS Bedrock(eu)
    claude-opus-4-8
    $5.50$27.50$0.55
    AWS Bedrock(us)
    claude-opus-4-8
    $5.50$27.50$0.55
    NovitaAI
    qwen3.7-max
    $1.25$3.75$0.25
    Granite
    qwen3.7-max
    $2.50$1.25
    -50% off
    $7.50$3.75
    -50% off
    $0.50$0.25
    -50% off
    Alibaba Cloud(us-virginia)
    qwen3.7-max
    $2.50$7.50$0.50
    Alibaba Cloud(singapore)
    qwen3.7-max
    $2.50$7.50$0.50
    Alibaba Cloud(cn-beijing)
    qwen3.7-max
    $1.72$5.17$0.34
    Alibaba Cloud
    qwen3.7-max
    $2.50$7.50$0.50
    Alibaba Cloud(eu-frankfurt)
    qwen3.7-max
    $1.65$4.95$0.33
    xAI
    grok-build-0-1
    $1.00$2.00$0.20
    Google Vertex AI
    gemini-3.5-flash
    $1.50$9.00$0.15
    Google AI Studio
    gemini-3.5-flash
    $1.50$9.00$0.15
    Vertex AI (OpenAI-compatible)
    grok-4-20-non-reasoning
    $1.25$2.50$0.20
    Vertex AI (OpenAI-compatible)
    grok-4-20-reasoning
    $1.25$2.50$0.20
    Xiaomi
    mimo-v2.5
    $0.14$0.28$0.00
    Xiaomi
    mimo-v2.5-pro
    $0.43$0.87$0.00
    Google AI Studio
    gemini-3.1-flash-lite
    $0.25$1.50$0.02
    Google Vertex AI
    gemini-3.1-flash-lite
    $0.25$1.50$0.02
    AWS Bedrock(global)
    grok-4-3
    $1.25$2.50$0.20
    AWS Bedrock(us)
    grok-4-3
    $1.38$2.75$0.22
    AWS Bedrock(us-west-2)
    grok-4-3
    $1.38$2.75$0.22
    AWS Bedrock
    grok-4-3
    $1.25$2.50$0.20
    xAI
    grok-4-3
    $1.25$2.50$0.20
    NovitaAI
    qwen3.6-35b-a3b
    $0.25$1.48—
    Alibaba Cloud(singapore)
    qwen3.6-35b-a3b
    $0.25$1.48—
    Alibaba Cloud(cn-beijing)
    qwen3.6-35b-a3b
    $0.25$1.48—
    Alibaba Cloud(eu-frankfurt)
    qwen3.6-35b-a3b
    $0.25$1.48—
    Alibaba Cloud
    qwen3.6-35b-a3b
    $0.25$1.48—
    Alibaba Cloud(us-virginia)
    qwen3.6-35b-a3b
    $0.25$1.48—
    Alibaba Cloud(eu-frankfurt)
    qwen3.6-plus
    $0.28$1.65$0.03
    Alibaba Cloud(cn-beijing)
    qwen3.6-plus
    $0.50$3.00$0.05
    Alibaba Cloud(singapore)
    qwen3.6-plus
    $0.50$3.00$0.05
    Alibaba Cloud
    qwen3.6-plus
    $0.50$3.00$0.05
    Page 3 of 9

    Newsletter

    Stay ahead of the curve

    Join developers who get weekly insights on LLM routing, new model launches, and cost optimization — straight to their inbox.

    • New models & providers as they drop
    • Tips to cut latency & costs
    • Early access to beta features

    No spam. Unsubscribe anytime.

    All systems operational
    AICPA SOC for Service Organizations badgeSOC 2 Type II
    compliant

    Product

    • Features
    • AI Gateway
    • Observability
    • Models
    • Providers
    • Rankings
    • Add Provider
    • Partners
    • Lounge
    • Changelog
    • DevPass
    • Compare Models
    • Enterprise

    Resources

    • Legal Overview
    • Apps
    • Templates
    • Agents
    • MCP Server
    • Use Cases
    • Blog
    • Documentation
    • Integrations
    • Guides
    • Brand Assets
    • Token Cost Calculator
    • Copilot Cost Calculator
    • Referral Program
    • GitHub
    • Discord
    • Twitter
    • Contact Us

    Compliance

    • Trust Center
    • Security Portal
    • Terms
    • Privacy Policy
    • Provider Information
    • Sub-processors
    • SOC 2 Type II
    • Status

    Compare

    • All Comparisons
    • GitHub Copilot
    • OpenRouter
    • LiteLLM
    • Portkey
    • AWS Bedrock
    • Azure AI Foundry
    • Vercel AI Gateway
    • Migration Guides

    Models

    • Text Generation
    • Text to Image
    • Image to Image
    • Video Generation
    • Embeddings
    • Vision
    • Reasoning
    • Tool Calling
    • Web Search
    • Discounted
    • Best for Roleplay
    • Best for Coding
    • Best for Creative Writing
    • Best for Translation
    • Best for Math
    • Long Context
    • Cheapest
    • Open Source

    Providers

    • OpenAI
    • Anthropic
    • Google AI Studio
    • Glacier
    • Iceberg
    • Granite
    • Google Vertex AI
    • Vertex AI (OpenAI-compatible)
    • Vertex AI (Anthropic)
    • Quartz
    • Avalanche
    • Groq
    • Cerebras
    • xAI
    • DeepSeek
    • Alibaba Cloud
    • NovitaAI
    • AtlasCloud
    • AWS Bedrock
    • AWS Mantle
    • Azure
    • Azure AI Foundry
    • Z AI
    • Moonshot AI
    • Baidu
    • Permafrost
    • Perplexity
    • Nebius AI
    • Mistral AI
    • CanopyWave
    • Inference.net
    • Together AI
    • SCX.ai (Turbo)
    • SCX.ai
    • Custom
    • NanoGPT
    • ByteDance
    • MiniMax
    • EmberCloud
    • Meta
    • Sakana AI
    • Tundra
    • Xiaomi
    • DeepInfra
    • Reve
    • ElevenLabs
    • Runware
    • Gonka24
    • Fireworks AI
    • RanoAI

    © 2026 LLM Gateway. All rights reserved.

    Every model on this page accepts at least 200,000 tokens of context — roughly 150,000 words — and the largest stretch much further: Grok 4.1 Fast at 2 million tokens, with Gemini, Claude Sonnet 5, GPT-5.4, DeepSeek V4, and GLM-5.2 at or above the million-token mark. That's enough to fit an entire codebase, a legal document set, or months of chat history into a single prompt.

    Advertised size isn't everything: retrieval quality can degrade well before the window is full, and long prompts get expensive fast. Cached input pricing — shown in the list — matters more than the headline price when you re-send large contexts on every request.

    Frequently asked questions

    Which LLM has the largest context window?

    Grok 4.1 Fast currently leads with a 2 million token window. Gemini models run just over 1 million, and Claude Sonnet 5, GPT-5.4, DeepSeek V4, GLM-5.2, and Qwen3.7 also offer million-token windows.

    How many words fit in a 200K context window?

    Roughly 150,000 English words — about 600 pages. A million-token window fits around 750,000 words: several full-length books, or a mid-sized codebase.

    Do models actually use the full window well?

    Not uniformly. Most models recall the start and end of a prompt better than the middle, and effective context is often smaller than the advertised maximum. For critical retrieval over huge inputs, test with your own data and consider chunking plus retrieval instead of one giant prompt.

    How do I keep long-context costs down?

    Use cached input pricing: providers charge a fraction of the normal rate for re-sent, unchanged prefixes, which is exactly the shape of chatting over a large document or codebase. Structure prompts so the big static context comes first and only the question changes.

    is now on LLM Gateway — 30% off open-source modelsends in 7d 10:25:01
    LLM Gateway
    • DevPass
    • Lounge
    • Models
    • Docs
    • Pricing
    • DevPass
    • Lounge
    • Pricing
    • Docs
    • Models
      • AI Gateway
      • DevPass
      • Lounge
      • Observability
      • Enterprise
      • Blog
      • Changelog
      • Integrations
      • Reliability
      • Guardrails
      • Providers
      • Partners
      • Rankings
      • Apps
      • Models
      • Model Timeline
      • Compare
      • Token Cost Calculator
      • Referral Program
      • MCP Server
      • Agents
      • AI SDK Provider
      • Agent Skills
      • Templates
      • Guides
    Log InGet Started
    1.6k