Prompt caching through OpenRouter
An agent thread re-sends its whole prefix on every model call: the tool definitions, the system prompt, and the history. A long bug thread does that tens of times per round and for dozens of rounds. Prompt caching lets the model provider keep that prefix and charge a fraction of the input price when the same bytes come again. This package does two things on every OpenRouter request so that the cache is written once and then read.
What is sent
| Field | Value | Why |
|---|---|---|
top-level cache_control |
{"type": "ephemeral", "ttl": "1h"} |
Providers that cache only on request (Anthropic, Google, Alibaba Qwen) cache the prefix up to the last cacheable block. OpenRouterPromptCachePolicy. |
top-level session_id |
the thread's id, mw- and 32 hex digits |
OpenRouter keeps every request with this id on the upstream provider that served the last one ("sticky routing"), so the provider that wrote the cache is the one that is read. A session expires after 10 minutes without a request, so the stickiness is shorter than the one-hour cache (see below). OpenRouterSessionWire. |
Both are added only for an OpenRouter endpoint (openrouter.ai or a subdomain, such as
eu.openrouter.ai). Another OpenAI-compatible gateway may reject an unknown top-level field, so it
gets neither. A request that already carries its own cache_control keeps it.
One hour, not five minutes
The default lifetime of a cached prefix is five minutes. Our threads rest longer than that between rounds: a bug-pool thread waits 20 minutes, and a person reads before answering. With five minutes every such round paid the whole history again. With one hour:
- a cache write costs 2× the input price instead of 1.25×, but only for the part that is new this round;
- every read within the hour costs about 0.1× and starts the hour again, so a thread worked at least hourly keeps its prefix.
The hour holds only on the provider that wrote the cache. The OpenRouter session that keeps a thread on that provider expires after 10 minutes without a request. A thread that rests longer, like the 20-minute bug-pool wait, is routed afresh on its next round. For a model served by one provider nothing changes. For a model our EU routing allows on several providers (Opus on Bedrock eu-west-1 and on Vertex europe), that round can land on the other provider. It then pays the full history at input price plus a fresh write there. So the saving above is certain only for rounds less than 10 minutes apart, or for single-provider models. Beyond that it depends on which provider the round lands on.
One session per thread
The id is derived from the thread's path (ProviderSession.For). It is the same after a restart, a
roll or a change of model, so nothing has to be stored for it to work. It is a hash, so no user
name, issue title or space name reaches the gateway. It is still recorded on the thread
(Thread.ProviderSessionId) the first time a round runs, so OpenRouter's activity log can be
matched to the thread.
The id travels from the round (AgentChatClient) to the client in the round's ChatOptions, under
ProviderSession.OptionsKey. The round puts it there only when its chat client answers
IProviderSessionClient, which only an OpenRouter session wire does. Any other provider
(Anthropic, Ollama, Copilot, a non-OpenRouter OpenAI endpoint) never sees the key. The wire still
takes it off the options before the request is built, so the internal key never reaches any wire.
Without the id OpenRouter falls back to a hash of the messages, which a growing thread changes every round. A model that our EU routing allows on several providers, such as Opus on Bedrock eu-west-1 and on Vertex europe, could then land each round on a different provider, whose cache holds nothing of the thread.
What each upstream provider does
| Provider | Caches | Write | Read |
|---|---|---|---|
| Anthropic (Bedrock, Vertex, Azure) | on cache_control |
1.25× (5 min) / 2× (1 h) | 0.1× |
| Google Gemini | automatically | input + storage | 0.25× |
| Alibaba Qwen | on cache_control |
1.25× | 0.1× |
| OpenAI | automatically | 1.25× on GPT-5.6+ | 0.25–0.5× |
| DeepSeek | automatically | 1× | 0.1× |
| Moonshot (Kimi), Z.AI (GLM), Grok | automatically | free | 0.2–0.25× |
The table is OpenRouter's documentation as read on 2026-10-03. Prefixes below a model's minimum (1,024 to 4,096 tokens) are not cached, without an error. Caches are separate per model, so a thread that moves to another model starts with an empty cache there.
What cannot be done
The cache is the model provider's, keyed by the exact bytes of the prompt. It cannot be read out, stored by us, or written back later. Sending the same prefix again is the only way to create or refresh an entry. A cache we own is possible only for a model we serve ourselves.
Not used yet
cache_controlon individual blocks: up to four breakpoints, for example a fixed one at the end of the shared system prompt besides the automatic one at the end of the history.- OpenAI's own
prompt_cache_options(mode and lifetime) andprompt_cache_breakpointon GPT-5.6 and later.
Is it working
The response reports cached_tokens and, where the provider writes, cache_write_tokens. The
engine reads them for the cost of a round (UsageTokens.SplitCache), but they are not yet stored
on the thread's messages or usage records. Until they are, the place to see cache reads per request
is the Activity page of the OpenRouter account.