Endpoint and authentication
The router exposes two OpenAI-compatible surfaces. Both share the same auth, routing, caching, and features.The
/responses endpoint is a translation layer over the same router:
identical auth, routing, caching, and features. Everything in this guide
applies to both surfaces unless a section says otherwise.SLNG_API_KEY, passed as a Bearer token. The
OpenAI SDK’s api_key parameter sets this header for you, so you never set
internal auth headers yourself.
Making a request
A request is a standard OpenAI Chat Completions call.Streaming
Streaming works exactly as in OpenAI, for every model the router serves (reasoning and non-reasoning alike).- Reasoning blocks and control tokens are removed for you. If your model emits
a visible reasoning block (
<think>,<thought>,<thinking>,<reasoning>) or a stray control token (</s>,<|im_end|>and similar), the router strips it before the chunk reaches you, so your TTS never reads it aloud. Settingslng_pure_proxysuspends this. - All answer content arrives before the
finish_reason: "stop"chunk, so stopping there is safe. - The
usageblock rides the final chunk. That is whyinclude_usagematters. - The stream always ends cleanly. If the model fails mid-stream, the router
still sends a final
finish_reason: "stop"chunk anddata: [DONE], so your client is never left hanging.
Available parameters
slng/auto vs a named model
slng/auto
Recommended. SLNG picks from your configuration: the preferred tier first,
splitting traffic by weight, with automatic retry on failure.
A named model
The exact name of one entry in your configuration. It pins the request to that
entry, so there is no failover to another one, and its traffic keeps its own
cache entries.
Required fields
Your org identity is managed by SLNG. The agent ID and session ID are yours to send, and both are required on every request.- Agent ID names the logical agent a call belongs to, and scopes its cache. Use
a stable value per agent. When you change that agent’s prompt in a meaningful way,
give it a fresh ID, for example
clinic-scheduler-v2, so answers produced under the old prompt are not reused. - Session ID identifies a single call. Use a value unique to each call, like a UUID, and keep it the same across all turns of that call.
X-Slng-Agent-Id, X-Slng-Session-Id) or as a
top-level body field (slng_agent_id, slng_session_id). The body field wins if
you send both. Values are at most 256 characters and cannot contain spaces, commas,
pipes (|) or braces ({ }). A request missing either ID, or carrying a
malformed one, returns a 400 with code missing_client_id or
invalid_client_id.
Reading the response
Every request returns a set ofx-slng-* headers that tell you how the answer
was produced.
On an error, only
x-slng-request-id comes back, because no answer was produced.
Two outcomes are possible on a successful request:
Normal LLM response
x-slng-response-source: llm. The request was routed to the LLM, and usage
carries the tokens you were billed for.Cache hit
x-slng-response-source: cache. No LLM ran. usage carries the original
answer’s tokens, so it is what the cache saved rather than what this call cost.Response caching
The Context Router is a PII-aware service that can serve a repeated turn straight from a cache instead of calling the LLM again. Caching is on by default, and a turn that qualifies is served from cache automatically. Cache layers are checked fastest-first; the first hit wins.Pre-warmed answers are set up together with you, per agent.
What SLNG does not store
So one caller’s answer is never replayed to another, some responses are not stored. The caller still gets their answer; only the caching is skipped.- Any reply containing a number, in any script. A number is usually specific to one caller or record, so identifiers, prices and dates are never cached.
- Personal information, such as names, addresses and emails.
- Any reply carrying a placeholder the router did not fill in itself.
Per-call values (template_variables)
Voice agents need per-call values inside the system prompt: the caller’s name,
their city, an appointment time. Instead of building those strings yourself
before every request, write a placeholder once and let SLNG fill it in.
Do this rather than building the string yourself. The stored answer then holds the
placeholder instead of the caller’s data, which is what makes a personalized
answer cacheable and lets your callers share one entry. The model sees the same
final prompt either way.
Syntax. Use double braces around the variable name:
{{caller_name}},
{{city}}. Names are letters, digits, and underscore. Single braces
({name}), Jinja, and ${var} are not placeholders. Only {{name}} is.
Substitution applies to your prompt messages, the system and developer roles.
Your user and assistant message content is never modified, so a literal
{{...}} a caller happens to type is left exactly as-is.You are calling Rajesh in Mumbai.
What belongs in template_variables
- Do carry personalization values that get spoken or echoed in the conversation: the customer’s name, the agent’s name, the company name, an appointment detail rendered as text.
- Don’t carry values that steer or change the content of the answer itself: the response language, a plan or tier that changes which policy is described, a region that changes which rules apply.
- The language case deserves spelling out. A language variable is fine when it is
just text to be said, for example the agent confirming “So your preferred
language is
{{language}}, correct?” during the call. It is not fine when it controls the language the model answers in: two callers who chose different languages would share one cached answer, and one would hear the wrong language.
Limits. Up to 64 variables per request, names up to 64 characters, values
up to 4000 characters. Exceeding these returns a
422.Runtime variables. To leave a placeholder deliberately unfilled, because your
own orchestrator fills it mid-call, set
template_vars_strict: false on your
configuration. The placeholder then passes through untouched, and an answer that
echoes one is never cached.BYOK Context Router (slng_config)
There are two ways the router can know which model(s) should answer your
requests:
Inline, per request
You send the configuration on each request. It runs against exactly what you
sent, and you change models by changing your own code.
Stored for your org
Register a BYOK LLM key in the dashboard and the router routes
slng/auto to
it with nothing extra in the request. Richer setups, such as several tiers or
failover groups, SLNG configures with you.A request carrying
slng_config ignores the stored configuration entirely, so
anything you rely on has to be in the object you send. It holds your own endpoints
and provider keys, so treat it like a credential.The configuration shape
A configuration is a set of numbered tiers, tried in order of preference:"1", "2", "3", three at most. Within a tier, traffic splits by weight.
If the chosen model fails to answer (a
5xx, a timeout, or a 429), the router
retries once against the next option. With a single entry it retries once
against that same endpoint when the failure was transient.
The smallest useful configuration is one tier with one OpenAI-compatible
endpoint:
Supported providers
Theendpoint object defaults to provider: "openai-compat", which covers any
service that speaks the OpenAI Chat Completions API (OpenAI itself, Azure-hosted
alternatives, Groq, self-hosted vLLM, and so on). Four more providers are
supported; each needs its own fields on the endpoint object.
On
azure, url is the resource root, not the full deployment URL. On
openai-compat, add auth_header when the provider wants its key in a custom
header. Fields belonging to another provider are rejected rather than ignored.
For example, an Azure OpenAI entry looks like:
Limits and errors
- The
slng_configobject must stay under 256 KB when serialized; larger returns a400. - A configuration that fails validation returns a
400starting withinvalid slng_config:and a description of the problem. Your credentials never appear in these messages. - Sending
slng_configas anything other than a JSON object (a string, a list) returns a422with codeinvalid_slng_config.
Shadow mode
Two optional flags let you route production traffic through SLNG and measure what it would save, before you let it change anything. Both default tofalse and need no
setup.
Send both together and you have a shadow trial. You receive exactly what your own
model produces, while the router measures what it could have served from cache and
what that would have saved. Going live is removing the two flags: same URL, same
key, same request body.
A trial measures; it does not speed anything up. A turn the router would have
answered from cache goes to your model instead, so trial latency reflects your own
model. Send your prompt in template form with
template_variables for the
measurement to be meaningful.