Skip to main content
A provider bills a cached prompt prefix at a discount, but only when every byte before the cache boundary matches the previous request. TrustGate sits in the middle of that prefix, so it does three things:
  • Same format, untouched. When your client and the provider speak the same API — an Anthropic client on an Anthropic registry, an OpenAI Chat client on OpenAI — the request goes upstream as you sent it. Cache markers are never added, moved or dropped. An OpenAI Chat client on Azure, DeepSeek, xAI, Cerebras or an OpenAI-compatible endpoint is same format too.
  • Cross format, translated. When they differ, your markers are rewritten into the target’s own mechanism, and anything the target does not accept is left out rather than sent.
  • Usage in your dialect. Cache reads and writes come back in the fields your client already reads, and they are priced and traced the same way for every provider.
TrustGate never inserts a cache marker of its own. If your client sends none, the only caching you get is what the provider does automatically.

What each provider needs

Some providers cache a repeated prefix automatically. Others cache only what the request marks. A provider decides for itself how long a prefix has to be before it caches it. A marker on a shorter prefix is accepted and ignored.

What each target receives

Your client marks cache boundaries in its own dialect: On a cross-format request, the target gets the part of that intent it accepts: Markers on assistant turns and images do not reach OpenAI Responses, and no target in the OpenRouter rows takes tool markers. When there are more markers than the target allows, they are trimmed the same way everywhere:
  1. At most 4 breakpoints reach Anthropic, Bedrock and OpenRouter. The earliest message breakpoints go first, then the earliest tool breakpoints. The last breakpoint of tools, system and messages is always kept.
  2. 1h comes before 5m. Providers require longer TTLs first, so a 1h breakpoint after a 5m one is sent as 5m. Where the target has no 1h cache, every breakpoint is sent as 5m.
  3. prompt_cache_options.mode: "explicit" is removed when no breakpoint is left, since explicit mode with nothing marked would turn caching off.
A same-format request skips all of this. Your markers reach the provider exactly as sent, and the provider’s own limits apply. An OpenAI Chat request routed to OpenRouter or Groq is the exception: it is re-encoded, so the rows above apply to it.
Routing decides which row you get. With fallback or load balancing, the same request can land on a provider that takes your markers and on one that does not. The request succeeds either way; only the cache hit rate differs.

Bedrock models

TrustGate sends cachePoint only to models that take it. The model is read after routing, with one inference-profile prefix (us., eu., apac., jp., au., ca., us-gov., global.) removed. A foundation-model or system inference-profile ARN resolves through the model ID it ends with. An application inference profile, provisioned throughput or custom model ARN does not say which model it runs, so it gets no cachePoint. If Bedrock still rejects a request over its cache checkpoints, TrustGate retries that request once without them, and the call is served uncached.

Examples

OpenAI Chat client to Claude on Anthropic

The client marks a long system prompt for an hour and the ticket history for five minutes. The question that changes on every call comes after the last marker.
Client request
Anthropic receives the same boundaries as Messages blocks. prompt_cache_key has no Anthropic equivalent and is left out.
What Anthropic receives (excerpt)
On the second call the client reads the hit where an OpenAI client always does:
Response usage (excerpt)

Anthropic client to Claude on Bedrock

The same markers on a Bedrock registry become cachePoint blocks, each placed right after what it covers.
Client request (excerpt)
What Bedrock receives (excerpt)
On Claude 3.7 Sonnet the ttl would be omitted, which is Bedrock’s 5m default. On Nova the tool cachePoint would be left out as well.

Usage and cost

A cached request has three token counts that matter: tokens read from the cache, tokens written to it, and the share of the writes made with a 1h TTL. TrustGate reports them in the fields your client already reads: Anthropic clients get cache_creation, with both the 5m and the 1h count, only when the provider reported that split. Streaming responses carry the same fields. Every trace records the same counts under usage, whichever dialect the client spoke: They are exported over OpenTelemetry as trustgate.usage.cached_input_tokens, trustgate.usage.cache_write_input_tokens and trustgate.usage.cache_write_1h_input_tokens. See the event schema. Cost. Each count is priced at its own rate: reads at the cache-read rate, writes at the cache-write rate, and the rest of the prompt at the input rate. A model with no published cache rate bills those tokens at its input rate. A 1h write costs twice the input rate on the Anthropic and Bedrock providers, which is what both charge for Claude, and the plain cache-write rate everywhere else. A registry’s contract pricing applies to cache rates too. The list discount reduces them with the input rate. A model override can set cache_read, cache_write and cache_write_1h through the registry API. An override that sets only input and output bills cached tokens at its input rate, and 1h writes on Anthropic and Bedrock at twice that. A token budget in LLM Budget does not count cache reads. A dollar budget counts them at the cache-read rate.

When caching is lost

Policies that edit the request

Most policies only judge a request. The ones that change it take one of two paths: Tool Injection and the Per-Tool Rate Limiter edit the request in place: in the common case, every byte they do not change goes upstream as you sent it. The masking policies rebuild it instead, so the body is no longer your exact bytes. It is still the same bytes on every request with the same input, so a mask that always hits the same text keeps the prefix cached. Policies that rewrite a buffered response keep the cache usage fields in it.

Tips

  • Put what does not change first. Tools, then the system prompt, then earlier turns. Anything that varies per call goes after the last marker.
  • Keep the system prompt stable. A timestamp, request ID or user name near the top invalidates everything after it. Prompt Template variables count too.
  • Send prompt_cache_key for OpenAI, Azure and Mistral. It is the one cache field all three take, and OpenAI Chat and Responses clients can send it to any of them. An Anthropic client has no field for it, so Mistral does not cache its requests. Use one key per prompt family, not one per user.
  • Mark the last stable block, not every block. Four breakpoints is the most any target takes, and extra ones are trimmed.
  • Choose 1h only for prefixes reused over the hour. A 1h write costs twice the input rate against the cheaper 5m write. It pays off when the prefix is read again after the five-minute cache would have expired.
  • Check the hit rate on the trace. No cached_input_tokens on a repeated request means the prefix changed, the target does not cache it, or the prefix is under the provider’s minimum.