Lightweight text embedding model with 1024-dimensional output and a 32K context, aimed at large-scale retrieval where latency and cost matter most.
LLM Gateway routes requests to the best providers that are able to handle your prompt size and parameters.
Lightweight text embedding model with 1024-dimensional output and a 32K context, aimed at large-scale retrieval where latency and cost matter most. You can access it through LLM Gateway's OpenAI-compatible API with automatic provider routing, fallback, and cost analytics.
Pricing for Kinfra Text Embedding 0.6B on LLM Gateway starts at $0.07 per million input tokens and $0.00 per million output tokens, depending on the provider. The pricing table above always reflects the current per-provider rates.
Kinfra Text Embedding 0.6B supports a context window of up to 32,768 tokens on its largest provider deployment.
Kinfra Text Embedding 0.6B is served by Tencent Cloud through LLM Gateway. Requests are automatically routed to the best available provider, with fallback when a provider has issues.
No. Kinfra Text Embedding 0.6B does not currently support tool calling or structured JSON outputs through LLM Gateway.
Kinfra Text Embedding 0.6B was released on June 30, 2026.