LLM API: Call 250+ Models Through One Endpoint
An LLM API tutorial: get a key, send your first chat completion to https://api.llmgateway.io/v1, switch models by changing one string, stream tokens, pin or let the gateway pick the provider, and keep using the OpenAI and Anthropic SDKs you already have.
By LLM Gateway

Every model provider ships its own LLM API: a different base URL, a different auth header, a different request body, and a different way of streaming. Supporting three of them means three clients, three sets of error codes and three places where an outage can take your product down. LLM Gateway replaces that with one OpenAI-compatible LLM API at https://api.llmgateway.io/v1: one key, one request shape, 250+ models from 40+ providers, and automatic fallback when a provider fails.
This is the shortest path from zero to a working call, then the four things you will want next: switching models, streaming, controlling which provider answers, and using the SDKs you already have.
Get an LLM API key
- Sign in at llmgateway.io and create a project.
- Copy the project's API key.
- Export it in your shell:
1export LLM_GATEWAY_API_KEY="llmgtwy_XXXXXXXXXXXXXXXX"1export LLM_GATEWAY_API_KEY="llmgtwy_XXXXXXXXXXXXXXXX"New accounts start with free credits. Bring your own provider keys and the gateway adds no markup on top of provider prices.
Send your first chat completion
1curl https://api.llmgateway.io/v1/chat/completions \2 -H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "model": "gpt-4o",6 "messages": [{"role": "user", "content": "Explain what an LLM API is in two sentences."}]7 }'1curl https://api.llmgateway.io/v1/chat/completions \2 -H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "model": "gpt-4o",6 "messages": [{"role": "user", "content": "Explain what an LLM API is in two sentences."}]7 }'The response is the standard OpenAI chat completion object: choices[0].message.content holds the text, usage holds token counts. The gateway adds a metadata object naming the provider, model and region that served the request, and listing every attempt under routing when a fallback happened.
Switch models by changing one string
The model field accepts any ID from the model catalog. Nothing else in the request changes:
| Want | model |
|---|---|
| Anthropic's current mid-tier | claude-sonnet-5-5 |
| Google's fast model | gemini-3.8-flash |
| OpenAI's coding model | gpt-6.1-sol |
| An open-weight model served by several providers | deepseek-v3.2 |
| A free model to test with | ling-3.1-flash |
List what your key can use programmatically:
1curl https://api.llmgateway.io/v1/models \2 -H "Authorization: Bearer $LLM_GATEWAY_API_KEY"1curl https://api.llmgateway.io/v1/models \2 -H "Authorization: Bearer $LLM_GATEWAY_API_KEY"Each entry includes context size, pricing per provider and supported parameters, so your app can pick a model at runtime instead of hard-coding one.
Stream tokens
Add "stream": true and read server-sent events. The chunk format is OpenAI's for every model, including providers whose native LLM API streams differently:
1curl https://api.llmgateway.io/v1/chat/completions \2 -H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "model": "claude-sonnet-5-5",6 "stream": true,7 "messages": [{"role": "user", "content": "Write a haiku about rate limits."}]8 }'1curl https://api.llmgateway.io/v1/chat/completions \2 -H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "model": "claude-sonnet-5-5",6 "stream": true,7 "messages": [{"role": "user", "content": "Write a haiku about rate limits."}]8 }'Usage and routing metadata arrive on the final chunk, so you can log cost per request without a second call.
Control which provider answers
A bare model ID such as deepseek-v3.2 can be served by several providers. By default the gateway scores every healthy one on price, uptime, throughput and time to first token, picks the best, and retries on another if the first fails. Two ways to steer that:
1# Bias toward the cheapest healthy provider2-d '{"model": "deepseek-v3.2", "routing": "price", "messages": [...]}'3
4# Bias toward the highest-throughput provider5-d '{"model": "deepseek-v3.2", "routing": "throughput", "messages": [...]}'1# Bias toward the cheapest healthy provider2-d '{"model": "deepseek-v3.2", "routing": "price", "messages": [...]}'3
4# Bias toward the highest-throughput provider5-d '{"model": "deepseek-v3.2", "routing": "throughput", "messages": [...]}'To pin one provider with no fallback, prefix it:
1-d '{"model": "openai/gpt-4o", "messages": [...]}'1-d '{"model": "openai/gpt-4o", "messages": [...]}'For multi-turn conversations, send an x-session-id header (or the OpenAI prompt_cache_key / user body fields). The gateway pins that session to one provider and region so the upstream prompt cache stays warm. Details are in the routing docs.
Keep the SDKs you already use
Because the LLM API is OpenAI-compatible, the official OpenAI SDKs work by changing the base URL:
1import OpenAI from "openai";2
3const client = new OpenAI({4 baseURL: "https://api.llmgateway.io/v1",5 apiKey: process.env.LLM_GATEWAY_API_KEY,6});7
8const completion = await client.chat.completions.create({9 model: "gemini-3.8-flash",10 messages: [{ role: "user", content: "Summarize this ticket in one line." }],11});1import OpenAI from "openai";2
3const client = new OpenAI({4 baseURL: "https://api.llmgateway.io/v1",5 apiKey: process.env.LLM_GATEWAY_API_KEY,6});7
8const completion = await client.chat.completions.create({9 model: "gemini-3.8-flash",10 messages: [{ role: "user", content: "Summarize this ticket in one line." }],11});1from openai import OpenAI2import os3
4client = OpenAI(5 base_url="https://api.llmgateway.io/v1",6 api_key=os.environ["LLM_GATEWAY_API_KEY"],7)8
9completion = client.chat.completions.create(10 model="claude-haiku-5-5",11 messages=[{"role": "user", "content": "Classify this email: spam or not?"}],12)13print(completion.choices[0].message.content)1from openai import OpenAI2import os3
4client = OpenAI(5 base_url="https://api.llmgateway.io/v1",6 api_key=os.environ["LLM_GATEWAY_API_KEY"],7)8
9completion = client.chat.completions.create(10 model="claude-haiku-5-5",11 messages=[{"role": "user", "content": "Classify this email: spam or not?"}],12)13print(completion.choices[0].message.content)Code written against the Anthropic SDK or Claude Code does not need to change either: the gateway also serves /v1/messages in Anthropic's format, and any catalog model can be called through it. Set ANTHROPIC_BASE_URL=https://api.llmgateway.io and ANTHROPIC_AUTH_TOKEN=$LLM_GATEWAY_API_KEY. See the Anthropic endpoint docs.
Vercel AI SDK users can use createOpenAI({ baseURL: "https://api.llmgateway.io/v1", apiKey: process.env.LLM_GATEWAY_API_KEY }) (the provider otherwise reads OPENAI_API_KEY) or the dedicated @llmgateway/ai-sdk-provider package.
Beyond chat: the same key for the rest of the LLM API
The same base URL and key cover /v1/embeddings, /v1/images/generations, /v1/videos, /v1/moderations, /v1/rerank, /v1/search and /v1/responses, all with the standard OpenAI error envelope. Gateway-enforced rate limits return 429 with error.code: "rate_limit_exceeded" and RateLimit-* headers. Upstream provider failures go through retry and fallback first; when the gateway wraps a non-OpenAI-shaped provider error, the error body carries requestedProvider and usedProvider; OpenAI-shaped upstream errors are forwarded unchanged and may not include those fields. See error handling and rate limits.
Why one LLM API instead of three
- One integration, every model. Adding a provider is a catalog change, not a code change.
- Fallback is on by default. A provider outage becomes a routing event in
metadata.routing, not a page. - Cost and latency are visible per request in the dashboard, broken down by model, project and key.
- No markup on your own keys. Platform credits are optional.
- Open source. Self-host the gateway if the hosted LLM API does not fit your compliance needs.
Next steps
- Try LLM Gateway free and send the
curlabove with your own key. - Read the quickstart for Go, Rust, Java, PHP and Ruby examples.
- Compare providers and prices for any model on the models directory, or read what an LLM gateway is if you are deciding whether you need one.
Frequently asked questions
- What is an LLM API?
- An LLM API is an HTTP interface for sending prompts to a large language model and receiving generated text, tool calls or structured output. Each provider (OpenAI, Anthropic, Google, Mistral and others) ships its own, with its own base URL, auth header and request shape. LLM Gateway exposes one OpenAI-compatible LLM API, https://api.llmgateway.io/v1, that reaches 250+ models from 40+ providers.
- Can I use the OpenAI SDK with a different model provider?
- Yes. Point the SDK's baseURL at https://api.llmgateway.io/v1, use your LLM Gateway key, and set model to any catalog ID such as claude-sonnet-5-5, gemini-3.8-flash or deepseek-v3.2. The request and response shapes stay the same, so the rest of your code does not change.
- How do I choose which provider serves my request?
- Leave the model ID bare (for example deepseek-v3.2) and LLM Gateway scores every healthy provider on price, uptime, throughput and latency and picks the best one, with fallback if it fails. Add routing: "price" or routing: "throughput" to bias that choice, or prefix the provider (openai/gpt-4o) to pin it with no fallback.
- Does streaming work the same way?
- Yes. Send stream: true and the gateway returns server-sent events in the OpenAI chunk format for every provider, including ones whose native API streams differently. Usage and the provider that served the request are attached to the final chunk.