Support

AI-powered help

Welcome!

Please introduce yourself before we start.

    Cheapest Models

    Models at or under $0.20 per million input tokens ($1.50 output) — for classification, extraction, and high-volume workloads

    Compare
    Use Case
    Capabilities
    Provider
    Status
    Input Price ($/M tokens)
    Output Price ($/M tokens)
    Context Size (tokens)
    73/258
    Models
    32/55
    Providers
    30
    Vision Models (filtered)
    50
    Tool-enabled (filtered)
    0
    Free Models (filtered)
    Features
    NovitaAI
    ling-3.0-flash
    $0.06$0.18$0.01
    DeepInfra
    ling-3.0-flash
    $0.06$0.18$0.01
    DeepInfra
    hy3
    $0.14$0.58$0.04
    NovitaAI
    hy3
    $0.14$0.58$0.04
    Alibaba Cloud
    qwen3.7-flash
    $0.03$0.13$0.01
    Alibaba Cloud(cn-beijing)
    qwen3.7-flash
    $0.03$0.13$0.01
    Alibaba Cloud(singapore)
    qwen3.7-flash
    $0.03$0.13$0.01
    Alibaba Cloud(us-virginia)
    qwen3.7-flash
    $0.03$0.13$0.01
    Alibaba Cloud(eu-frankfurt)
    qwen3.7-flash
    $0.03$0.13$0.01
    Alibaba Cloud
    qwen-audio-3.0-tts-flash
    $0.015/1K chars——
    Alibaba Cloud
    qwen-audio-3.0-tts-plus
    $0.02/1K chars——
    xAI
    grok-stt-1-0
    $0.00$0.00—
    Nebius AI
    cosmos3-super-reasoner
    $0.10$0.30—
    Nebius AI
    nemotron-3-nano-omni
    $0.06$0.24—
    Nebius AI
    nemotron-3-nano-30b
    $0.06$0.24—
    Nebius AI
    hermes-4-70b
    $0.13$0.40—
    OpenAI
    gpt-5.6-luna
    $0.20$1.20$0.02
    Azure
    gpt-5.6-luna
    $0.20$1.20$0.02
    Alibaba Cloud(cn-beijing)
    qwen3.6-flash
    $0.17$0.99$0.03
    Alibaba Cloud(us-virginia)
    qwen3.6-flash
    $0.17$0.99$0.03
    Mistral AI
    mistral-ocr-latest
    $0.00$0.00—
    DeepInfra
    qwen3.5-9b
    $0.10$0.15—
    NovitaAI
    gemma-4-26b-a4b-it
    $0.13$0.40—
    DeepInfra
    gemma-4-26b-a4b-it
    $0.07$0.34—
    DeepInfra
    gemma-4-31b-it
    $0.13$0.38—
    Runware
    gemma-4-31b-it
    $0.10$0.07
    -30% off
    $0.30$0.21
    -30% off
    $0.01$0.01
    -30% off
    NovitaAI
    gemma-4-31b-it
    $0.14$0.40—
    ElevenLabs
    eleven-turbo-v2-5
    $0.055/1K chars——
    ElevenLabs
    eleven-flash-v2-5
    $0.055/1K chars——
    ElevenLabs
    eleven-v3
    $0.11/1K chars——
    ElevenLabs
    eleven-multilingual-v2
    $0.11/1K chars——
    OpenAI
    tts-1-hd
    $0.03/1K chars——
    OpenAI
    tts-1
    $0.015/1K chars——
    Xiaomi
    mimo-v2.5
    $0.14$0.28$0.00
    CanopyWave
    deepseek-v4-flash
    $0.14$0.28$0.03
    Together AI
    deepseek-v4-flash
    $0.14$0.28$0.03
    NovitaAI
    deepseek-v4-flash
    $0.14$0.28$0.03
    Alibaba Cloud(singapore)
    deepseek-v4-flash
    $0.20$0.40$0.04
    Runware
    deepseek-v4-flash
    $0.08$0.05
    -30% off
    $0.15$0.11
    -30% off
    $0.01$0.01
    -30% off
    ByteDance
    deepseek-v4-flash
    $0.14$0.28$0.03
    Gonka24
    deepseek-v4-flash
    $0.05$0.09$0.00
    Baidu
    deepseek-v4-flash
    $0.14$0.28$0.03
    DeepSeek
    deepseek-v4-flash
    $0.14$0.28$0.00
    Alibaba Cloud
    deepseek-v4-flash
    $0.20$0.40$0.04
    Alibaba Cloud(cn-beijing)
    deepseek-v4-flash
    $0.14$0.28$0.03
    Fireworks AI
    deepseek-v4-flash
    $0.14$0.28$0.03
    Alibaba Cloud(eu-frankfurt)
    deepseek-v4-flash
    $0.20$0.40$0.04
    Alibaba Cloud(us-virginia)
    deepseek-v4-flash
    $0.20$0.40$0.04
    DeepInfra
    deepseek-v4-flash
    $0.08$0.18$0.02
    EmberCloud
    qwen3-coder-next
    $0.11$0.68$0.06
    Page 1 of 3

    Newsletter

    Stay ahead of the curve

    Join developers who get weekly insights on LLM routing, new model launches, and cost optimization — straight to their inbox.

    • New models & providers as they drop
    • Tips to cut latency & costs
    • Early access to beta features

    No spam. Unsubscribe anytime.

    All systems operational
    AICPA SOC for Service Organizations badgeSOC 2 Type II
    compliant

    Product

    • Features
    • AI Gateway
    • Observability
    • Models
    • Providers
    • Rankings
    • Add Provider
    • Partners
    • Lounge
    • Changelog
    • DevPass
    • Compare Models
    • Enterprise

    Resources

    • Legal Overview
    • Apps
    • Templates
    • Agents
    • MCP Server
    • Use Cases
    • Blog
    • Documentation
    • Integrations
    • Guides
    • Brand Assets
    • Token Cost Calculator
    • Copilot Cost Calculator
    • Referral Program
    • GitHub
    • Discord
    • Twitter
    • Contact Us

    Compliance

    • Trust Center
    • Security Portal
    • Terms
    • Privacy Policy
    • Provider Information
    • Sub-processors
    • SOC 2 Type II
    • Status

    Compare

    • All Comparisons
    • GitHub Copilot
    • OpenRouter
    • LiteLLM
    • Portkey
    • AWS Bedrock
    • Azure AI Foundry
    • Vercel AI Gateway
    • Migration Guides

    Models

    • Text Generation
    • Text to Image
    • Image to Image
    • Video Generation
    • Embeddings
    • Vision
    • Reasoning
    • Tool Calling
    • Web Search
    • Discounted
    • Best for Roleplay
    • Best for Coding
    • Best for Creative Writing
    • Best for Translation
    • Best for Math
    • Long Context
    • Cheapest
    • Open Source

    Providers

    • OpenAI
    • Anthropic
    • Google AI Studio
    • Glacier
    • Iceberg
    • Granite
    • Google Vertex AI
    • Vertex AI (OpenAI-compatible)
    • Vertex AI (Anthropic)
    • Quartz
    • Avalanche
    • Groq
    • Cerebras
    • xAI
    • DeepSeek
    • Alibaba Cloud
    • NovitaAI
    • AtlasCloud
    • AWS Bedrock
    • AWS Mantle
    • Azure
    • Azure AI Foundry
    • Z AI
    • Moonshot AI
    • Baidu
    • Permafrost
    • Perplexity
    • Nebius AI
    • Mistral AI
    • CanopyWave
    • Inference.net
    • Together AI
    • SCX.ai (Turbo)
    • SCX.ai
    • Custom
    • NanoGPT
    • ByteDance
    • MiniMax
    • EmberCloud
    • Meta
    • Sakana AI
    • Tundra
    • Xiaomi
    • DeepInfra
    • Reve
    • ElevenLabs
    • Runware
    • Gonka24
    • Fireworks AI
    • RanoAI

    © 2026 LLM Gateway. All rights reserved.

    Every model here costs at most $0.20 per million input tokens and $1.50 per million output tokens — some, like Qwen3 4B and Llama 3.2 3B, as little as $0.03. At these prices a million-token workload costs a few cents, which changes what's economical: classify every support ticket, summarize every call, run an LLM check on every commit.

    Cheap doesn't mean toy: GPT-OSS 120B, GLM-4.7 Flash, Gemini 2.5 Flash-Lite, and Qwen3.5 9B punch far above their price on everyday tasks. Route high-volume work here and reserve frontier models for the requests that actually need them.

    Frequently asked questions

    What is the cheapest LLM API?

    In this catalog, Llama 3.2 3B and Qwen3 4B start around $0.03 per million input tokens, with GPT-OSS 20B at about $0.04 and GLM-4.7 Flash at $0.06. Prices differ per provider, so check the list — the same model is often cheaper through one provider than another.

    Are cheap models good enough for production?

    For classification, extraction, routing, summarization, and simple chat — usually yes. Small models fail mostly on multi-step reasoning and niche knowledge. A common pattern is a cheap model as the default with automatic escalation to a frontier model when confidence is low.

    How else can I cut LLM costs?

    Cache responses for repeated requests, use cached input pricing for long shared prefixes, batch offline work, and set per-project spending limits. Routing through a gateway also lets you switch to whichever provider currently offers the lowest price for the same model with zero code changes.

    Are there free models?

    Yes — free mappings show up at $0.00 in this list. They're rate-limited and best for prototyping; for production traffic, the paid models on this page are the reliable low-cost option.

    is now on LLM Gateway — 30% off open-source modelsends in 7d 12:15:45
    LLM Gateway
    • DevPass
    • Lounge
    • Models
    • Docs
    • Pricing
    • DevPass
    • Lounge
    • Pricing
    • Docs
    • Models
      • AI Gateway
      • DevPass
      • Lounge
      • Observability
      • Enterprise
      • Blog
      • Changelog
      • Integrations
      • Reliability
      • Guardrails
      • Providers
      • Partners
      • Rankings
      • Apps
      • Models
      • Model Timeline
      • Compare
      • Token Cost Calculator
      • Referral Program
      • MCP Server
      • Agents
      • AI SDK Provider
      • Agent Skills
      • Templates
      • Guides
    Log InGet Started
    1.6k