- Same format, untouched. When your client and the provider speak the same API — an Anthropic client on an Anthropic registry, an OpenAI Chat client on OpenAI — the request goes upstream as you sent it. Cache markers are never added, moved or dropped. An OpenAI Chat client on Azure, DeepSeek, xAI, Cerebras or an OpenAI-compatible endpoint is same format too.
- Cross format, translated. When they differ, your markers are rewritten into the target’s own mechanism, and anything the target does not accept is left out rather than sent.
- Usage in your dialect. Cache reads and writes come back in the fields your client already reads, and they are priced and traced the same way for every provider.
What each provider needs
Some providers cache a repeated prefix automatically. Others cache only what the request marks.
A provider decides for itself how long a prefix has to be before it caches it.
A marker on a shorter prefix is accepted and ignored.
What each target receives
Your client marks cache boundaries in its own dialect:
On a cross-format request, the target gets the part of that intent it accepts:
Markers on assistant turns and images do not reach OpenAI Responses, and no
target in the OpenRouter rows takes tool markers.
When there are more markers than the target allows, they are trimmed the
same way everywhere:
- At most 4 breakpoints reach Anthropic, Bedrock and OpenRouter. The earliest message breakpoints go first, then the earliest tool breakpoints. The last breakpoint of tools, system and messages is always kept.
- 1h comes before 5m. Providers require longer TTLs first, so a 1h breakpoint after a 5m one is sent as 5m. Where the target has no 1h cache, every breakpoint is sent as 5m.
prompt_cache_options.mode: "explicit"is removed when no breakpoint is left, since explicit mode with nothing marked would turn caching off.
Routing decides which row you get. With fallback or load
balancing, the same request can land on a provider that
takes your markers and on one that does not. The request succeeds either way;
only the cache hit rate differs.
Bedrock models
TrustGate sendscachePoint only to models that take it. The model is read
after routing, with one inference-profile prefix (us., eu., apac., jp.,
au., ca., us-gov., global.) removed.
A foundation-model or system inference-profile ARN resolves through the model
ID it ends with. An application inference profile, provisioned throughput or
custom model ARN does not say which model it runs, so it gets no
cachePoint.
If Bedrock still rejects a request over its cache checkpoints, TrustGate
retries that request once without them, and the call is served uncached.
Examples
OpenAI Chat client to Claude on Anthropic
The client marks a long system prompt for an hour and the ticket history for five minutes. The question that changes on every call comes after the last marker.Client request
prompt_cache_key
has no Anthropic equivalent and is left out.
What Anthropic receives (excerpt)
Response usage (excerpt)
Anthropic client to Claude on Bedrock
The same markers on a Bedrock registry becomecachePoint blocks, each placed
right after what it covers.
Client request (excerpt)
What Bedrock receives (excerpt)
ttl would be omitted, which is Bedrock’s 5m default.
On Nova the tool cachePoint would be left out as well.
Usage and cost
A cached request has three token counts that matter: tokens read from the cache, tokens written to it, and the share of the writes made with a 1h TTL. TrustGate reports them in the fields your client already reads:
Anthropic clients get
cache_creation, with both the 5m and the 1h count, only
when the provider reported that split. Streaming responses carry the same
fields.
Every trace records the same counts under usage, whichever dialect the
client spoke:
They are exported over OpenTelemetry as
trustgate.usage.cached_input_tokens,
trustgate.usage.cache_write_input_tokens and
trustgate.usage.cache_write_1h_input_tokens. See the event
schema.
Cost. Each count is priced at its own rate: reads at the cache-read rate,
writes at the cache-write rate, and the rest of the prompt at the input rate. A
model with no published cache rate bills those tokens at its input rate. A 1h
write costs twice the input rate on the Anthropic and Bedrock providers,
which is what both charge for Claude, and the plain cache-write rate everywhere
else.
A registry’s contract pricing
applies to cache rates too. The list discount reduces them with the input rate.
A model override can set cache_read, cache_write and cache_write_1h
through the registry API. An override that sets only input and output bills
cached tokens at its input rate, and 1h writes on Anthropic and Bedrock at
twice that.
A token budget in LLM Budget does not count
cache reads. A dollar budget counts them at the cache-read rate.
When caching is lost
Policies that edit the request
Most policies only judge a request. The ones that change it take one of two paths:
Tool Injection and the Per-Tool Rate Limiter edit the request in place: in the
common case, every byte they do not change goes upstream as you sent it. The
masking policies rebuild it instead, so the body is no longer your exact bytes.
It is still the same bytes on every request with the same input, so a mask that
always hits the same text keeps the prefix cached.
Policies that rewrite a buffered response keep the cache usage fields in it.
Tips
- Put what does not change first. Tools, then the system prompt, then earlier turns. Anything that varies per call goes after the last marker.
- Keep the system prompt stable. A timestamp, request ID or user name near the top invalidates everything after it. Prompt Template variables count too.
- Send
prompt_cache_keyfor OpenAI, Azure and Mistral. It is the one cache field all three take, and OpenAI Chat and Responses clients can send it to any of them. An Anthropic client has no field for it, so Mistral does not cache its requests. Use one key per prompt family, not one per user. - Mark the last stable block, not every block. Four breakpoints is the most any target takes, and extra ones are trimmed.
- Choose 1h only for prefixes reused over the hour. A 1h write costs twice the input rate against the cheaper 5m write. It pays off when the prefix is read again after the five-minute cache would have expired.
- Check the hit rate on the trace. No
cached_input_tokenson a repeated request means the prefix changed, the target does not cache it, or the prefix is under the provider’s minimum.