Real-time reliability for every provider serving Text Embedding 004 on LLM Gateway. Compare token volume, success rates, time-to-first-token, throughput, and error breakdown across 1 provider so you can pick the fastest, most stable route for your workload.
Each card shows hourly traffic from the last 24 hours. Switch tabs to inspect tokens, requests, errors, or latency.
Every Text Embedding 004 request flowing through LLM Gateway is scored on uptime, latency, and throughput. When an upstream provider degrades, traffic shifts to the next-best healthy endpoint without any client-side changes.
Use this page to verify SLA performance, debug regressions, or pick a primary provider for a self-hosted deployment.
Uptime is the share of valid requests that completed successfully over the last 24 hours. Client errors (4xx from your request) are excluded so the number reflects service reliability, not invalid requests.
Text Embedding 004 is currently served by 1 provider: Google Vertex AI. LLM Gateway routes requests to the best healthy provider in real time.
TTFT (time to first token) is the latency between the request and the first streamed token. Lower TTFT means the model starts responding faster — critical for chat UIs and agent loops.
Charts refresh every minute and aggregate the most recent 24 hours of traffic across all LLM Gateway projects, bucketed by hour. The current hour fills in as traffic arrives.