In early April 2026, Google released Gemma 4. I pulled it on Ollama, pointed my MTGAI pipeline at it, and watched a 26-billion-parameter model generate Magic: The Gathering cards on my desktop PC. A machine with a single consumer GPU and 12 gigs of video memo...
In the intro to this series, we built a naive 26-billion-parameter model and calculated that it would need 400 GB of GPU memory. That's five datacenter GPUs just to hold it in memory.
Let's start shrinking it.
The first two optimizations are the oldest and m...
We're at 202 GB. Last time we cut the model in half by switching to FP16. The model weights are down to 52 GB, but the KV cache - the model's memory of the conversation - is still 147 GB at 256k context. That's the bottleneck now.
We're at 112 GB. The KV cache is down to 58 GB thanks to GQA, but it's still the biggest single cost. Every layer in the model is storing attention notes for every token in the 256,000-token context. All 30 layers, all 256k tokens.
We're at 60 GB. Three optimizations - FP16, GQA, and sliding window attention - crushed the KV cache from 294 GB to under 6 GB. The problem has flipped: nearly all of the r...
We're at 60 GB. The KV cache is solved. MoE makes inference fast enough for real-time use and offloading survivable. But 52 GB of model weights are still sitting there, stored in FP16 - 65,536 possible value...
We're at 19 GB. Down from 400 GB. The major optimizations are done - FP16, GQA, sliding window, MoE, and quantization got us from "needs a datacenter" to "fits on a consumer GPU."
Gemma 4 adds two more innovations of its own. Neither is as dramatic as wh...