Perplexity Search API on LLM Gateway
Use the Perplexity Search API through LLM Gateway to get ranked web results with extracted page content for agents and retrieval pipelines, billed per search with the same key you use for every model.

An agent that answers questions about current events needs fresh web results,
and most of the time it needs the results themselves, not another model's
summary of them. Until now, the only way to reach Perplexity through LLM
Gateway was Sonar, which searches and then writes prose. The Perplexity
Search API is now available at /v1/search: ranked web results with extracted
page content, on the same key, bill and logs as every model you already call.
Raw search results or a search-grounded answer
Both approaches exist on the gateway, and they solve different problems.
| You want | Use |
|---|---|
| A model to search and answer in one call | Native web search |
| Results to rank, filter, cache or cite yourself | Search API (/v1/search) |
| Search as a tool for any model, including open models | Search API, called from your tool handler |
| Full control over which pages reach the prompt | Search API |
Native web search is the shortest path to an answer, but the model decides what to search and what to keep. The Search API hands that decision back to you: you get the URLs and snippets, and you choose what goes into the context window.
Call the Perplexity Search API
The endpoint mirrors Perplexity's request and response, so existing Perplexity Search code only needs a new base URL and key.
1curl https://api.llmgateway.io/v1/search \2 -H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "query": "latest developments in open-source LLMs",6 "max_results": 5,7 "search_recency_filter": "week",8 "search_domain_filter": ["arxiv.org", "huggingface.co"]9 }'1curl https://api.llmgateway.io/v1/search \2 -H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "query": "latest developments in open-source LLMs",6 "max_results": 5,7 "search_recency_filter": "week",8 "search_domain_filter": ["arxiv.org", "huggingface.co"]9 }'1{2 "id": "9ed15dce-a498-40b2-8bc9-7f13f35901cd",3 "model": "perplexity/perplexity-search",4 "results": [5 {6 "title": "…",7 "url": "https://…",8 "snippet": "…",9 "date": "2026-09-25",10 "last_updated": "2026-09-27"11 }12 ],13 "server_time": null14}1{2 "id": "9ed15dce-a498-40b2-8bc9-7f13f35901cd",3 "model": "perplexity/perplexity-search",4 "results": [5 {6 "title": "…",7 "url": "https://…",8 "snippet": "…",9 "date": "2026-09-25",10 "last_updated": "2026-09-27"11 }12 ],13 "server_time": null14}Every Perplexity filter is forwarded unchanged: country, language, domain
allow and deny lists, recency, publication and last-updated date ranges, and
per-result or total token budgets for extracted content. Pass query as an
array of up to five related queries to run them in a single request.
Give any model a search tool
Because the results are plain JSON, the Search API works as a tool for any model on the gateway. The handler below runs a search and returns compact results for the model to read and cite.
1async function searchWeb(query: string) {2 const res = await fetch("https://api.llmgateway.io/v1/search", {3 method: "POST",4 headers: {5 Authorization: `Bearer ${process.env.LLM_GATEWAY_API_KEY}`,6 "Content-Type": "application/json",7 },8 body: JSON.stringify({9 query,10 search_type: "fast",11 max_results: 5,12 }),13 });14 const { results } = await res.json();15 return results.map((r: { title: string; url: string; snippet: string }) => ({16 title: r.title,17 url: r.url,18 snippet: r.snippet,19 }));20}1async function searchWeb(query: string) {2 const res = await fetch("https://api.llmgateway.io/v1/search", {3 method: "POST",4 headers: {5 Authorization: `Bearer ${process.env.LLM_GATEWAY_API_KEY}`,6 "Content-Type": "application/json",7 },8 body: JSON.stringify({9 query,10 search_type: "fast",11 max_results: 5,12 }),13 });14 const { results } = await res.json();15 return results.map((r: { title: string; url: string; snippet: string }) => ({16 title: r.title,17 url: r.url,18 snippet: r.snippet,19 }));20}Register searchWeb as a function tool in the tools array of a
/v1/chat/completions request and return its output as the tool result. The
search call and the model call both land in the same activity log, so you can
see what each answer was built from.
Pricing: two models, billed per search
model is optional. Leave it out and Perplexity's search_type selects the
model; set it explicitly to pin one.
search_type | Model | Price |
|---|---|---|
web (or omitted) | perplexity/perplexity-search | $5 per 1,000 searches |
fast | perplexity/perplexity-search-fast | $1 per 1,000 searches |
Billing follows Perplexity's own rules:
- No token charges, however much page content a search returns.
- One charge per request, even when
queryholds several queries. - Failed requests are free. An upstream error is logged with a cost of zero.
Search requests count toward the same credits, spend limits, IAM rules and data
retention settings as the rest of your traffic. Perplexity's people search
type is not supported yet and returns a 400.
Get started
Frequently asked questions
- What is the difference between the Perplexity Search API and Sonar?
- Sonar is a chat model that searches the web and writes an answer. The Search API returns the ranked results themselves, with titles, URLs, snippets and dates, and leaves the reasoning to your own model or code.
- How much does the Perplexity Search API cost on LLM Gateway?
- perplexity/perplexity-search costs $5 per 1,000 searches and perplexity/perplexity-search-fast costs $1 per 1,000 searches. There are no token charges, a request with several queries counts as one search, and failed requests are not billed.
- Do I need to change my Perplexity Search code?
- No. The /v1/search endpoint accepts the same request body and returns the same response shape as Perplexity's Search API. Point your client at https://api.llmgateway.io/v1/search and use your LLM Gateway API key.