Support

AI-powered help

Welcome!

Please introduce yourself before we start.

    Open Source Models

    Open-weight models — Llama, DeepSeek, Qwen, GLM, Kimi, GPT-OSS, Gemma, and more — served through one API

    Compare
    Use Case
    Capabilities
    Provider
    Status
    Input Price ($/M tokens)
    Output Price ($/M tokens)
    Context Size (tokens)
    76/258
    Models
    28/55
    Providers
    25
    Vision Models (filtered)
    62
    Tool-enabled (filtered)
    0
    Free Models (filtered)
    Features
    NovitaAI
    ernie-4.5-vl-424b-a47b
    $0.42$1.25—
    Fireworks AI
    kimi-k3-fast
    $4.50$22.50$0.45
    Nebius AI
    minicpm-v-4.5
    $0.66$1.11—
    Nebius AI
    cosmos3-super-reasoner
    $0.10$0.30—
    Nebius AI
    nemotron-3-nano-omni
    $0.06$0.24—
    Nebius AI
    nemotron-3-nano-30b
    $0.06$0.24—
    Nebius AI
    nemotron-3-super-120b
    $0.30$0.90—
    Nebius AI
    hermes-4-70b
    $0.13$0.40—
    Nebius AI
    hermes-4-405b
    $1.00$3.00—
    Alibaba Cloud(singapore)
    glm-5.2
    $1.40$4.40$0.28
    CanopyWave
    glm-5.2
    $1.40$4.40$0.26
    NovitaAI
    glm-5.2
    $1.40$4.40$0.26
    Alibaba Cloud(us-virginia)
    glm-5.2
    $1.40$4.40$0.28
    EmberCloud
    glm-5.2
    $1.26$3.96$0.23
    Z AI
    glm-5.2
    $1.40$4.40$0.26
    Alibaba Cloud
    glm-5.2
    $1.40$4.40$0.28
    Baidu
    glm-5.2
    $1.40$4.40$0.26
    Alibaba Cloud(cn-beijing)
    glm-5.2
    $1.40$4.40$0.28
    Runware
    glm-5.2
    $0.80$0.56
    -30% off
    $2.55$1.79
    -30% off
    $0.16$0.11
    -30% off
    Alibaba Cloud(eu-frankfurt)
    glm-5.2
    $1.40$4.40$0.28
    ByteDance
    glm-5.2
    $1.40$4.40$0.26
    SCX.ai
    glm-5.2
    $0.55$1.78$0.11
    Nebius AI
    glm-5.2
    $1.40$4.40—
    DeepInfra
    qwen3.5-9b
    $0.10$0.15—
    Moonshot AI
    kimi-k2.7-code-highspeed
    $1.90$8.00$0.38
    NovitaAI
    gemma-4-26b-a4b-it
    $0.13$0.40—
    DeepInfra
    gemma-4-26b-a4b-it
    $0.07$0.34—
    Cerebras
    gemma-4-31b-it
    $0.99$1.49—
    Together AI
    gemma-4-31b-it
    $0.39$0.97—
    DeepInfra
    gemma-4-31b-it
    $0.13$0.38—
    Runware
    gemma-4-31b-it
    $0.10$0.07
    -30% off
    $0.30$0.21
    -30% off
    $0.01$0.01
    -30% off
    SCX.ai (Turbo)
    gemma-4-31b-it
    $0.30$0.91—
    NovitaAI
    gemma-4-31b-it
    $0.14$0.40—
    Nebius AI
    kimi-k2.7-code
    $0.95$4.00—
    NovitaAI
    kimi-k2.7-code
    $0.95$4.00$0.19
    Moonshot AI
    kimi-k2.7-code
    $0.95$4.00$0.19
    Nebius AI
    nemotron-3-ultra-550b
    $1.00$3.00—
    DeepInfra
    nemotron-3-ultra-550b
    $0.50$2.20$0.10
    Nebius AI
    minimax-m3
    $0.30$1.20—
    Together AI
    minimax-m3
    $0.30$1.20$0.06
    MiniMax
    minimax-m3
    $0.60$2.40$0.12
    Xiaomi
    mimo-v2.5
    $0.14$0.28$0.00
    Xiaomi
    mimo-v2.5-pro
    $0.43$0.87$0.00
    NovitaAI
    qwen3.6-35b-a3b
    $0.25$1.48—
    Alibaba Cloud(singapore)
    qwen3.6-35b-a3b
    $0.25$1.48—
    Alibaba Cloud(cn-beijing)
    qwen3.6-35b-a3b
    $0.25$1.48—
    Alibaba Cloud(eu-frankfurt)
    qwen3.6-35b-a3b
    $0.25$1.48—
    Alibaba Cloud
    qwen3.6-35b-a3b
    $0.25$1.48—
    Alibaba Cloud(us-virginia)
    qwen3.6-35b-a3b
    $0.25$1.48—
    CanopyWave
    deepseek-v4-flash
    $0.14$0.28$0.03
    Page 1 of 5

    Newsletter

    Stay ahead of the curve

    Join developers who get weekly insights on LLM routing, new model launches, and cost optimization — straight to their inbox.

    • New models & providers as they drop
    • Tips to cut latency & costs
    • Early access to beta features

    No spam. Unsubscribe anytime.

    All systems operational
    AICPA SOC for Service Organizations badgeSOC 2 Type II
    compliant

    Product

    • Features
    • AI Gateway
    • Observability
    • Models
    • Providers
    • Rankings
    • Add Provider
    • Partners
    • Lounge
    • Changelog
    • DevPass
    • Compare Models
    • Enterprise

    Resources

    • Legal Overview
    • Apps
    • Templates
    • Agents
    • MCP Server
    • Use Cases
    • Blog
    • Documentation
    • Integrations
    • Guides
    • Brand Assets
    • Token Cost Calculator
    • Copilot Cost Calculator
    • Referral Program
    • GitHub
    • Discord
    • Twitter
    • Contact Us

    Compliance

    • Trust Center
    • Security Portal
    • Terms
    • Privacy Policy
    • Provider Information
    • Sub-processors
    • SOC 2 Type II
    • Status

    Compare

    • All Comparisons
    • GitHub Copilot
    • OpenRouter
    • LiteLLM
    • Portkey
    • AWS Bedrock
    • Azure AI Foundry
    • Vercel AI Gateway
    • Migration Guides

    Models

    • Text Generation
    • Text to Image
    • Image to Image
    • Video Generation
    • Embeddings
    • Vision
    • Reasoning
    • Tool Calling
    • Web Search
    • Discounted
    • Best for Roleplay
    • Best for Coding
    • Best for Creative Writing
    • Best for Translation
    • Best for Math
    • Long Context
    • Cheapest
    • Open Source

    Providers

    • OpenAI
    • Anthropic
    • Google AI Studio
    • Glacier
    • Iceberg
    • Granite
    • Google Vertex AI
    • Vertex AI (OpenAI-compatible)
    • Vertex AI (Anthropic)
    • Quartz
    • Avalanche
    • Groq
    • Cerebras
    • xAI
    • DeepSeek
    • Alibaba Cloud
    • NovitaAI
    • AtlasCloud
    • AWS Bedrock
    • AWS Mantle
    • Azure
    • Azure AI Foundry
    • Z AI
    • Moonshot AI
    • Baidu
    • Permafrost
    • Perplexity
    • Nebius AI
    • Mistral AI
    • CanopyWave
    • Inference.net
    • Together AI
    • SCX.ai (Turbo)
    • SCX.ai
    • Custom
    • NanoGPT
    • ByteDance
    • MiniMax
    • EmberCloud
    • Meta
    • Sakana AI
    • Tundra
    • Xiaomi
    • DeepInfra
    • Reve
    • ElevenLabs
    • Runware
    • Gonka24
    • Fireworks AI
    • RanoAI

    © 2026 LLM Gateway. All rights reserved.

    Open-weight models have closed most of the gap with proprietary frontiers: DeepSeek V4, Qwen3.7, GLM-5, Kimi K2, and MiniMax M3 sit near the top of real-world leaderboards, joined by OpenAI's GPT-OSS and Google's Gemma releases. Their weights are public — but running a 200B+ parameter model yourself means serious GPU infrastructure.

    This page lists open-weight models served by hosted providers, so you get the openness — inspectable weights, no lock-in, the option to self-host later — with API convenience. LLM Gateway itself is open source (AGPLv3) and self-hostable, so the whole stack can run on your terms.

    Frequently asked questions

    What is the best open source LLM?

    DeepSeek V4, Qwen3.7, GLM-5.2, Kimi K2.6, and MiniMax M3 are the current leaders, each within striking distance of proprietary frontier models. For smaller, hardware-friendly options, GPT-OSS 20B, Gemma 4, and Qwen3.5 9B are the standouts.

    What does 'open source' mean for LLMs?

    Usually 'open weight': the trained weights are downloadable, but licenses vary — some are Apache 2.0 or MIT, others (like the Llama license) carry usage restrictions, and training data is rarely published. Check the license of a specific model before building on it.

    Should I self-host or use an API?

    Self-hosting pays off with steady high volume, strict data-residency needs, or fine-tuned weights. For everything else, per-token APIs are cheaper than idle GPUs. A middle path: develop against hosted open models and keep self-hosting as an exit option, since the weights are public.

    Are open models cheaper than proprietary ones?

    Dramatically, per token. Competition among hosts drives prices down — DeepSeek V4 Flash and Qwen3 Coder 30B cost 10–50x less than frontier proprietary models. The list above shows every provider's price for each model.

    is now on LLM Gateway — 30% off open-source modelsends in 7d 12:16:12
    LLM Gateway
    • DevPass
    • Lounge
    • Models
    • Docs
    • Pricing
    • DevPass
    • Lounge
    • Pricing
    • Docs
    • Models
      • AI Gateway
      • DevPass
      • Lounge
      • Observability
      • Enterprise
      • Blog
      • Changelog
      • Integrations
      • Reliability
      • Guardrails
      • Providers
      • Partners
      • Rankings
      • Apps
      • Models
      • Model Timeline
      • Compare
      • Token Cost Calculator
      • Referral Program
      • MCP Server
      • Agents
      • AI SDK Provider
      • Agent Skills
      • Templates
      • Guides
    Log InGet Started
    1.6k