Some of your cache writes should have been reads. Here is how to find them.
A cache write is avoidable when the same bytes already had a live entry in the same scope, and each one pays 1.25x where a 0.1x read would have done. The test, the causes, and how to find yours.
Split an agent's cache traffic into writes and reads, as in the write-cost post, and one column deserves a second look. Some of those writes are writes that should have been reads: the content they wrote was byte-identical to content that already had a live cache entry at the moment they wrote it.
That is not "the cache was cold" or "the prefix changed". It is the narrow, checkable case where everything was in place for a read — same bytes, entry alive, same scope — and the call paid the 1.25x write premium anyway. Every one of those writes is avoidable spend, and an agent can carry a steady stream of them while every standard metric says it is caching fine.
What "avoidable" means, precisely
Calling a write avoidable is a strong claim, so here is the exact test, applied to each write on its own. A write counts as avoidable only when all three hold:
- A prior call had written the same prefix. Same
prefix_hash— the fingerprint the proxy computes over the serialized system prompt and tools — so the bytes were identical, not merely similar. - That entry was still alive. The write landed inside the TTL window of the earlier one. A write 20 minutes after a 5-minute entry expired is a cold-cache write, not an avoidable one, however much it looks like a repeat.
- Same cache scope. Same account, same platform, same model. A cache entry warmed on one scope is invisible to another, and a write from a different scope is unavoidable by definition.
Anything failing any test was excluded. What survives is the uncomfortable category: the provider had the bytes, the entry was warm, and the request wrote anyway.
How a warm entry gets re-written
If the bytes match and the entry is alive, why would the API write instead of read? The causes we traced all reduce to the same shape: the request did not present the prefix the way the cache stored it.
The breakpoint moved. A cache entry is keyed by the span up to the cache_control marker, not by the request as a whole. A caller that marks a different block on each call — the whole system prompt on one, only its first paragraph on the next — presents a different cacheable span each time, and a span the cache has not seen is a write, even though the underlying text never changed. We see this most in hand-rolled caching code that computes "where to put the marker" from something that varies, like message count.
The prefix was byte-identical but the span was not. Tools serialize ahead of the system prompt inside the cached span. Two calls with the same system prompt and the same tools in a different order have the same content and different bytes — the case the tool-ordering post covers. When the order flaps between two stable arrangements, the two variants keep re-writing each other's entries: A writes, B writes, A writes again, every entry alive and none ever read.
Two writers racing. Two instances of the same agent starting within the same second both miss, both write. The second write is avoidable in accounting terms, unavoidable in practice — you would need coordination to stop it, and one duplicated write per cold start is cheap.
The first two causes are worth engineering time. The third is worth knowing about so you do not chase it.
What an avoidable write costs
The arithmetic is the write-cost post's, applied to a write that never needed to happen. An avoidable write pays 1.25x the input rate on its prefix where a read would have paid 0.1x — on that span, twelve and a half times the price of the read it displaced. A 10,000-token prefix at $3 per million input tokens costs $0.0375 to write and $0.003 to read, so each avoidable write throws away about three and a half cents, every time it happens. On a fleet, the count is what matters: how many of an agent's writes had a warm entry sitting right next to them.
The detection matters more than the dollar figure. An agent with a high avoidable-write count looks healthy on every standard metric — it caches, it has entries, its hit rate is non-zero — and it is quietly running at a fraction of its possible saving. Nothing in a provider bill or a default dashboard distinguishes an avoidable write from a necessary one. You have to keep per-prefix history and check each write against it, which is exactly what the proxy's prefix fingerprinting exists to make possible: every call carries a prefix_hash, so "did a live entry exist for these bytes" is a query, not a guess.
What to do
Log the cache-write and cache-read token counts per call if you are not already, and alongside them a hash of the serialized system prompt plus tools — a cheap FNV-1a over the concatenation is enough. Then look for the signature of avoidable writing: the same hash writing repeatedly inside one TTL window. If you find it, check the two usual suspects in order — a cache breakpoint whose position is computed from something that varies per call, and a tool array whose serialization order is not stable. Both fixes are small, and both convert every subsequent write on that prefix into a read at a tenth of the price.