Back to blog

LiteLLM vs OpenRouter vs LLM Gateway: Which Proxy for Production?

LiteLLM vs OpenRouter vs LLM Gateway compared on the decisions that matter in production: who runs the proxy, what you pay on your own keys, how routing and failover work, what you get for observability, and how the three score on an independent latency benchmark.

By LLM Gateway

A glowing three-way junction on the central chip of a dark circuit board where one light trace splits toward three smaller chips, surrounded by balance scale, server rack, cloud, gauge and coin-stack icons, representing a LiteLLM vs OpenRouter vs LLM Gateway comparison

Three names come up in every "which LLM proxy" thread: LiteLLM, the Python proxy you self-host; OpenRouter, the hosted catalog with one key for everything; and LLM Gateway, the open-source gateway you can run yourself or use hosted. They solve the same problem, one OpenAI-compatible endpoint in front of many providers, with different trade-offs on who operates it, what it costs on your own keys, and how much it does for you once a provider fails.

This comparison sticks to the decisions that matter once traffic is real. Figures are from each product's public pricing and docs as of October 11, 2026; LLM Gateway numbers are from the model catalog and docs.

LiteLLM vs OpenRouter vs LLM Gateway at a glance

Decision LiteLLM OpenRouter LLM Gateway
What it is Python SDK + self-hosted proxy Hosted API Open-source gateway, hosted or self-hosted
Who runs it You OpenRouter LLM Gateway (managed) or you (Docker)
Open source Yes No Yes (AGPLv3)
Models 100+ providers via SDK 400+ models, 80+ providers 250+ models, 40+ providers
Bring your own keys Yes, no fee $25k/month of list-price inference free (Standard and Business), then 5%; Enterprise custom Yes, 0% fee
Platform credits None Fee on purchases: 5.5% Standard, 8% Business; no token markup 5% platform fee
Routing Manual fallback lists per model Provider-based routing Weighted scoring on uptime, throughput, price, latency
Failover Configured per model Yes Automatic, up to 2 retries, listed in response metadata
Response caching Built in, needs config Beta Built in, 10 s to 1 year TTL
Guardrails External integrations Enterprise Enterprise (prompt injection, PII, jailbreak, secrets)
Analytics Dashboard + spend logs Dashboard, logs, OpenTelemetry export Per-request cost, latency, cache hits; 90-day audit logs
Image and video generation Image yes, video limited Image yes, video limited Image and video (/v1/images/generations, /v1/videos)
Independent latency benchmark Not measured 88.3 composite (5th of 6) 90.8 composite (1st of 6)

Who operates the proxy

This is the first fork in the road.

LiteLLM is software. You deploy the proxy, give it a database for keys and spend, put it behind a load balancer, upgrade it, and page someone when it falls over. In return you own everything: the config, the logs, the network path. Teams with a platform group and strict data-residency rules often prefer this.

OpenRouter is a service. There is nothing to run; you get one key and a catalog of 400+ models. You also get nothing to configure beyond the request, and no option to keep traffic inside your own network.

LLM Gateway does both. The hosted endpoint at https://api.llmgateway.io/v1 works like OpenRouter: one key, no infrastructure. The same code is open source under AGPLv3 and runs with one Docker command when you need it inside your VPC. You can start hosted and move to self-hosted later without changing the API your application calls. See self-hosting LLM Gateway.

What you pay on your own keys

If you already have provider accounts, the question is what the proxy charges on top.

  • LiteLLM: nothing. The software is free; you pay for compute, the database and the people.
  • OpenRouter: on the Standard and Business plans, BYOK traffic is free up to $25,000 of list-price inference a month, then 5%, measured against what the same requests would cost at OpenRouter list prices. Credit purchases carry a 5.5% fee on Standard and 8% on Business; token prices themselves pass through without markup. Enterprise terms are custom.
  • LLM Gateway: 0% on your own keys, hosted or self-hosted. If you use platform credits instead, a 5% platform fee applies.

For a team whose BYOK traffic amounts to $40,000 a month of list-price-equivalent inference, that is $0 on LiteLLM and LLM Gateway, and $750 a month on OpenRouter Standard or Business for the $15,000 above the allowance.

Routing and failover when a provider fails

Open-weight models such as deepseek-v3.2 or kimi-k3 are served by several providers at different prices and uptime. How the proxy picks one, and what it does when the pick fails, decides your error rate.

LiteLLM routes with fallback lists you write per model. It works, and it is explicit, but the list does not know that one provider has been returning 529s for ten minutes unless you add health checks and tune them.

OpenRouter routes by provider preference and reroutes on failure.

LLM Gateway scores every healthy provider for the requested model on uptime, throughput, price and time to first token, sends to the best, and retries on another when the first fails. The response's metadata.routing lists every attempt, so a failover is visible per request instead of inferred from a spike in latency. Send "routing": "price" or "routing": "throughput" to bias the choice, or prefix a provider (deepinfra/deepseek-v3.2) to pin it. For multi-turn work, an x-session-id header pins the session to one provider so the upstream prompt cache stays warm. Details in the routing docs.

Observability

LiteLLM ships a dashboard and spend logs per virtual key; audit logs are an Enterprise feature. OpenRouter has improved its request logs, exports and OpenTelemetry support, all inside its own dashboard. LLM Gateway logs every request with the provider that served it, cost, latency and cache hit, broken down by project and key, keeps 90 days of audit logs, and exposes the same data over the API for your own dashboards; see enterprise LLM analytics.

Latency: the independent numbers

In the August 7, 2026 run of computesdk's open AI gateway benchmark, LLM Gateway ranked first of six gateways with a composite score of 90.8, driven by the tightest tail latency in the field (warm TTFT p95 of 1,151 ms, cold end-to-end p95 of 1,262 ms). OpenRouter scored 88.3, fifth. LiteLLM was not in the run, which measures hosted gateways. The full breakdown and how to rerun it are in our benchmark write-up.

Migrating between them

All three are OpenAI-compatible, so migration is a base URL change and, for provider-prefixed IDs, a model string change:

1import OpenAI from "openai";2
3const client = new OpenAI({4  baseURL: "https://api.llmgateway.io/v1",5  apiKey: process.env.LLM_GATEWAY_API_KEY,6});7
8// LiteLLM "gemini/gemini-3.8-flash" or OpenRouter "google/gemini-3.8-flash"9// become a bare catalog ID; add a provider prefix only to pin it.10const completion = await client.chat.completions.create({11  model: "gemini-3.8-flash",12  messages: [{ role: "user", content: "Hello" }],13});

Per-product details are in LLM Gateway vs LiteLLM and LLM Gateway vs OpenRouter.

Verdict

  • Pick LiteLLM if you have a platform team that wants to own the proxy end to end, traffic must never leave your network, and you are comfortable writing and maintaining fallback config.
  • Pick OpenRouter if you want the largest hosted catalog, will never self-host, and the credit fee (5.5% Standard, 8% Business) or the 5% BYOK fee above the $25k monthly allowance is acceptable.
  • Pick LLM Gateway if you want hosted convenience today with a self-hosting exit, 0% on your own keys, scored routing with visible failover, and the lowest tail latency on the one public benchmark that compares gateways.

Next steps

Frequently asked questions

Is LiteLLM a gateway or an SDK?
Both. LiteLLM is a Python SDK that normalizes 100+ provider APIs behind an OpenAI-style call, and a proxy server you self-host to expose that as an OpenAI-compatible endpoint with virtual keys and budgets. There is no hosted LiteLLM endpoint; you run and scale it yourself.
What is the difference between OpenRouter and LiteLLM?
OpenRouter is a hosted service: one API key, 400+ models, nothing to run, with a fee on credit purchases (5.5% on Standard, 8% on Business) and a 5% fee on bring-your-own-key traffic above a $25,000-per-month allowance of list-price inference; Enterprise terms are custom. LiteLLM is open-source software you deploy and operate; it charges nothing but you pay for the infrastructure and the on-call.
When should I pick LLM Gateway over LiteLLM or OpenRouter?
When you want the hosted convenience of OpenRouter with the self-hosting option and zero-markup BYOK of LiteLLM in one product. LLM Gateway is open source (AGPLv3), runs as a managed service or one Docker command, charges 0% on your own keys, and scores providers on uptime, throughput, price and latency instead of static fallback lists.
Which LLM proxy adds the least latency?
In computesdk's independent AI gateway benchmark (August 7, 2026 run), LLM Gateway ranked first of six gateways with a composite score of 90.8 and a warm p95 time-to-first-token of 1,151 ms, ahead of OpenRouter at 88.3. LiteLLM was not in that run; as self-hosted software its overhead depends on where you deploy it.