The Cache Key I Only Sent to OpenAI
2026-09-13
This whole piece rests on three requests. I sent them to api.x.ai, all with the same prompt prefix, and wrote the numbers into the commit message of 99112ec. Here they are before anything else, because everything I'm going to say is just me reading this table.
| Probe | prompt_cache_key | Conversation | Cached / prompt tokens | Cached |
|---|---|---|---|---|
| 1 | sent | cold | 512 / 5758 | 8.9% |
| 2 | sent | warm | 5632 / 5796 | 97.2% |
| 3 | not sent | warm | 512 / 5771 | 8.9% |
Rows 2 and 3 are the same warm conversation. The only difference is one field in the request body. With it, 97.2% of the prompt came out of cache. Without it, 8.9%, which is exactly what a cold request gets.
From Wizard 2.0 on 2026-08-10 until that commit, every Grok turn Wizard made went out like row 3, with no key.
Reading row 3 against row 1
Row 3 is the one I keep looking at. Without the key, the warm conversation got 512 cached tokens. The keyed cold request in row 1 also got 512. So the missing key didn't switch caching off. Something got cached either way. What the key does is get the request back to the server that already holds the prefix.
That matters because xAI's cache is per server. Their docs, as quoted in src/llm/wire.rs (I checked them on 2026-09-13), describe prompt_cache_key as "a stable cache key for best-effort sticky routing / prompt-cache hits across requests sharing a prompt prefix", plumbed to x-grok-conv-id. No key, no stickiness. A turn lands on whatever server answers, and that server has usually never seen the conversation. The grok-4.6 page, as the commit quotes it, puts it plainly: without it "you often pay full input price on a cache-cold server".
So row 3 isn't a broken cache. It's a working cache on the wrong machine.
Reading row 2 against row 3 in dollars
The easy headline is "one missing field cost 4x." That's not what the table says. The 4 comes from somewhere else.
On grok-4.6, fresh input is $2.00 per million tokens and cached input is $0.50. Cached reads cost a quarter. That's the 4x: a per-token price ratio, and it's the ceiling on what a miss can cost you, not a measurement.
The measurement is the table. Working it out in per-million units:
| Cached × $0.50 | Fresh × $2.00 | Input cost | |
|---|---|---|---|
| Row 2, keyed | 5632 × 0.50 | 164 × 2.00 | 3144 |
| Row 3, no key | 512 × 0.50 | 5259 × 2.00 | 10774 |
That's 3.43x (python3 -c "print((512*.5+5259*2)/(5632*.5+164*2))" gives 3.4268). At this probe's size it's about $0.0108 a request against $0.0031. It's under 4x because even a perfect hit leaves a few fresh tokens at the tail, and even a miss got 512 for free.
What I can't give you is the total. These are single requests. Nothing on the box recorded cache rate across whole Grok sessions before and after the fix, so I don't have a number for how much input I overpaid between 08-10 and 09-13, and I'm not going to invent one. The doc comment on with_cache_key in src/plugins/xai.rs says why it still matters: at those rates, "on a long agent turn this is most of the input bill."
How it lasted 34 days
Wizard 2.0 (9e0c054) added prompt_cache_key gated on is_openai_api(base_url), which matched one constant, OPENAI_API_ROOT = "https://api.openai.com". The doc comment explained the gate: the key "is a field of this endpoint's Chat Completions API, not of the wire shape, so it is matched on rather than guessed at."
That reasoning was correct and the conclusion was too narrow. Sending unknown fields to every OpenAI-compatible server is a bad idea. But xAI speaks the same wire shape and documents the same field, and the gate had no way to know that.
Then on 2026-08-24 I split the OpenAI wire protocol out of the OpenAI provider (1f4a164) and moved the remaining eight providers into plugins (be73eb5). The gate came along unchanged, still OpenAI-only. A refactor is exactly when you'd expect someone to reread a condition like that, and it's exactly when nobody does, because the job is to move code without changing what it does. It didn't change what it did. 3.1.0 shipped on 2026-09-11 with it.
The fix
99112ec replaced the single constant with a list:
CACHE_KEY_API_ROOTS = &["https://api.openai.com", "https://api.x.ai"]
endpoint_takes_a_cache_key in wire.rs matches against it anchored, so https://api.x.ai.evil.test doesn't count. Every xAI route gets the key: kind = "xai", kind = "xaioauth", and an xAI base URL set under kind = "openai". The test for it is a_grok_turn_carries_the_cache_key_and_a_relay_does_not in src/plugins/xai.rs.
I kept the original instinct. vLLM, LM Studio, llama.cpp, DeepSeek and Cloudflare still don't get it, since there it's "a field on the wire that is at best ignored and at worst rejected by a strict server."
The key itself is wz- plus eight bytes of a SHA-256 over the leading run of system messages and the model tag. System notes that show up mid-history, like the tool-failure nudge or a subagent report, are left out, so a failed tool call doesn't re-key the session and throw it onto a fresh server. The commit is 4 files, +425 -238. In docs/usage.md I also raised the top of the xAI cached-rate range from 0.2x to 0.25x, because grok-4.6 reads at 0.25x and the old top end was wrong.
The other half, 24 seconds later
23d928c landed at 16:57:35, right after the key fix at 16:57:11. It's the same bug from the other side. Routing the request to the warm server does nothing if I then edit the history it cached.
Wizard used to shrink old tool results once there was 16k characters to reclaim. But the provider's cached prefix ends at the first edit. Touching the oldest result in a 150k-token history made the next request re-prefill about 145k tokens at full price, about 22 cents on grok-4.6, to reclaim 4k tokens that saved half a cent a step. A hundred steps to break even. Now a shrink pass also has to reclaim at least a tenth of what its earliest edit invalidates (RECLAIM_TAIL_DIVISOR = 10 in src/agent/context.rs). xAI's own advice is "never modify earlier messages, only append new ones". The 60-step session test didn't move: 33315 to 18174 prompt tokens per call on average, 63484 to 18062 at the last step.
Both went out in 3.2.0 at 18:36 that evening, under Fixed: "xAI requests carried no prompt cache key, so almost nothing was cached."