We're at 202 GB. Last time we cut the model in half by switching to FP16. The model weights are down to 52 GB, but the KV cache - the model's memory of the conversation - is still 147 GB at 256k context. That's the bottleneck now.
To understand why, and how to fix it, we need to talk about attention heads.
16 Perspectives, 16 Copies
Recall that our model pushes every token through 30 layers - 30 processing stages stacked on top of each other, each one refining the model's understanding. Early layers work close to the raw words; late layers think in abstract concepts. And each of those 30 layers has 16 attention heads - 16 different lenses on the same text. One head might learn to track subject-verb relationships. Another might notice when a pronoun refers to something mentioned earlier. A third might pick up on tone or sentiment. They run in parallel, each producing its own Keys and Values (as we saw with the horse sentence last time), and their results get combined.
Here's the expensive part: each of those 16 heads maintains its own Key and Value vectors for every token in the context. That's 16 complete copies of the conversation's "notes" per layer. Across 30 layers and 256,000 tokens, it adds up to 147 GB.
But do all 16 heads really need their own private notes? Look at what our three heads stored for "horse" in the previous post:
| Head | Key | Value |
|---|---|---|
| Grammar | noun, subject, animal | large animal, the actor |
| Meaning | living creature, large mammal | strength, speed, four-legged |
| Narrative | main character, agent | who the story is about |
Three different sets of notes. But look at the overlap. All three Keys say some version of "important noun, the thing doing stuff." All three Values say some version of "it's the horse." Do we really need to store this three times - let alone sixteen?
The Spectrum of Sharing
In 2019, a researcher named Noam Shazeer (then at Google) proposed an extreme answer: Multi-Query Attention (MQA). What if all 16 query heads shared a single set of Keys and Values?
Picture a classroom. In the original setup (Multi-Head Attention), every student has their own personal copy of the textbook. With MQA, there's one copy on the teacher's desk and all 16 students reference it.
The memory savings are dramatic. Instead of 16 KV heads per layer, you have 1:
FP16 with 16 KV heads: ~0.57 MB per token
FP16 with 1 KV head: ~0.04 MB per token
At 256k context, that KV cache drops from 147 GB to about 9 GB. But there's a catch - quality degrades noticeably. All 16 heads looking at identical Keys and Values means less diverse attention patterns. The model becomes less nuanced, like trying to understand a novel from 16 identical perspectives instead of 16 different ones.
MQA went too far.
The Goldilocks Solution
Four years later, in 2023, a team at Google Research found the middle ground: Grouped Query Attention (GQA). Instead of every head having its own K/V (too expensive) or all heads sharing one K/V (too lossy), group the heads. Put a few students per table and give each table its own copy of the textbook. The paper was a direct refinement of Shazeer's MQA - built by his former colleagues at Google while he was away at Character.AI, improving the idea he'd started.
Meta's Llama 2 70B, released that July, was the first major model to ship with GQA, using 8 KV heads shared across 64 query heads - groups of 8. The quality difference from full MHA was minimal. The memory savings were massive.
Gemma 4 uses GQA with different groupings depending on the layer type (we'll explain the local/global distinction in the next post - for now, just know that Gemma 4 has two kinds of layers with different attention ranges):
- Local attention layers (short-range): 2 query heads per KV head (8 KV heads)
- Global attention layers (full-context): 8 query heads per KV head (2 KV heads)
The KV cache drops dramatically:
FP16 with full MHA (16 KV heads): ~0.57 MB/token --> 147 GB at 256k
FP16 with GQA (8/2 KV heads): ~0.23 MB/token --> 58 GB at 256k
89 GB gone, just by sharing textbooks. There's something deeply satisfying about this one - it's not clever math or exotic hardware, just noticing that everyone was making too many photocopies.
Why Does Sharing Work?
This is the part that wasn't obvious. If each head is supposed to learn a different perspective, wouldn't forcing heads to share Keys and Values cripple their ability to specialize?
Come back to the horse. Instead of three separate K/V pairs, GQA stores one shared set:
| Token | Shared Key | Shared Value |
|---|---|---|
| horse | primary actor | powerful animal in motion |
Two words, two words. But watch what happens when different heads query against it:
- Grammar head's Query: "what's the grammatical subject?" - primary actor → the subject is an animal
- Meaning head's Query: "what kind of creature?" - primary actor → a powerful living thing
- Narrative head's Query: "who's the protagonist?" - primary actor → the central character, currently in motion
Same Key, same Value, three different reads. This works because the three original Keys - "noun," "creature," "agent" - aren't as different as they look. They live near each other in vector space: all three describe "an important thing that does stuff." A single vector like "primary actor" can sit in that neighborhood and serve all three heads well enough. The heads don't need their own Keys to specialize - their Queries do that work. A shared textbook that each head highlights differently.
Sidebar: The man who invented attention sharing
Multi-Query Attention was invented in 2019 by Noam Shazeer at Google. Shazeer was also a co-author of "Attention Is All You Need," the 2017 paper that introduced the transformer architecture underneath every model in this series - so he was optimizing his own invention.
In 2021, Shazeer left Google, reportedly frustrated that the company wouldn't publicly release the chatbot products built on his research. CEO Sundar Pichai personally asked him to stay. He left anyway and co-founded Character.AI.
Three years later, Google paid $2.7 billion - roughly a dollar per parameter at GPT-3 scale - to bring him back. By then, the attention-sharing techniques Shazeer had pioneered were standard across the industry, and Google had finally shipped those chatbot products as Bard and then Gemini. Google essentially paid a massive premium to re-acquire talent and ideas it had in-house three years earlier. Shazeer is now VP of Engineering and co-lead of Gemini.
The Running Total
The memory breakdown is shifting. When we started, the KV cache was the overwhelming cost - 294 GB out of 400 GB. After FP16, it was still the majority at 147 GB out of 202 GB. Now with GQA, it's getting closer to the weight cost:

| Step | Weights | KV Cache | Overhead | Total |
|---|---|---|---|---|
| Naive (FP32) | 104 GB | 294 GB | 2.6 GB | ~400 GB |
| + FP16 | 52 GB | 147 GB | 2.6 GB | ~202 GB |
| + GQA | 52 GB | ~58 GB | 2.6 GB | ~112 GB |
One hundred and twelve gigabytes. Down from 400. We're not on a consumer GPU yet, but we're within striking distance of a single high-end datacenter GPU. And we haven't even touched the model weights - all of this reduction has come from the KV cache.
Why This Is the Context Window Story
The KV cache is the reason context windows exist. Not as a design choice - as a memory constraint.
GPT-3 (2020) had a context window of 2,048 tokens - about 1,500 words. GPT-3.5 launched at 4,096. GPT-4 started at 8,192. These weren't arbitrary limits. They were the largest contexts the hardware could afford to cache, given the KV memory cost at the time.
Every optimization in this series - GQA, sliding window attention, quantization - didn't just make models smaller. It made longer conversations possible. The jump from GPT-3's 2k context to Gemini's 1-million-token window isn't just better hardware. It's the KV cache going from 0.57 MB per token (full MHA, FP16) to a fraction of that through the techniques we're walking through.
This also explains something practical: when you run a model locally, the context size you configure directly determines how much memory it uses. The model weights are fixed - Gemma 4's 18 GB download is 18 GB regardless. But the KV cache scales with context. Set your context to 4k tokens and it barely registers. Set it to 256k and it might not fit.
That linear scaling is itself an achievement. Computing attention used to require memory that scaled with the square of context length - double the context, quadruple the memory. Flash Attention (Dao et al., 2022) fixed this by reorganizing the computation so the full attention matrix never needs to exist in memory at once. That's a story for another post, but it's a big part of why 256k contexts are practical at all.
Which brings me to an apology I owe Ollama. In my last post, I complained about Ollama silently limiting my context window below what the model supports. Having now spent several thousand words explaining why the KV cache is an enormous memory hog that scales linearly with context length... yeah, I get it. Defaulting to a conservative context window on consumer hardware is a reasonable call. A warning would've been nice, though.
The next optimization is going to take another enormous bite out of that remaining 58 GB, by asking a simple question: does every layer really need to see the entire conversation?
Next up: The Thousand-Token Window: Why Most Layers Don't Need to See Everything | Series start: How 400 GB Became 18
Written by Claude as dictated by Harald.
Comments (0)
No comments yet. Be the first!