Gemma 4 Gets Multi-Token Prediction Drafters: 3x Faster Inference Without Quality Loss

Canonical version: Gemma 4 Gets Multi-Token Prediction Drafters: 3x Faster Inference Without Quality Loss.

On 2026-05-05, Google released a companion line of small autoregressive drafter models for the Gemma 4 family, plus a Multi-Token Prediction (MTP) head. Available on Hugging Face and Kaggle. Supported in Transformers, MLX, vLLM, SGLang, and Ollama from day one. Apache 2.0.

Google reports up to 3× decoding speedup with no quality degradation, and around 2.2× on Apple Silicon with mixture-of-experts variants and batch sizes 4–8.

Through 2024–2025, Speculative Decoding was a generic add-on: you paired a small model with a larger one. Gemma 4's drafters are co-designed with the target model.

The drafters rely on three architectural changes:

  • Target activation sharing. The drafter consumes the final-layer activations of the target Gemma 4 model on round 1, concatenated with its own embeddings. The target's prompt encoding is reused instead of recomputed.
  • KV cache sharing. The drafter cross-attends to the target's KV cache instead of building its own. This keeps the drafter small without losing long-context information.
  • Efficient embedder. Gemma 4 has a 262K-token vocabulary. The drafter uses sparse decoding via clustered token lookup; identify the most likely token cluster, then compute logits only inside it. This is two-stage retrieval applied to the LM head.

For the deeper architectural breakdown, see AI Multi-Token Prediction Drafters.

Consequences of the release:

  1. Gemma 4 is the first major open-weight family to ship a drafter alongside the main model. Other labs will likely do the same.
  2. The gains are largest in memory-bandwidth-bound setups: single user, batch size 1, consumer GPU or Apple Silicon, which is how local LLMs typically run. Datacenter GPUs already amortize memory transfers across batched users; consumer hardware doesn't.
  3. The target model isn't retrained. The drafters are independent artifacts. Existing Gemma 4 weights stay valid; you just add the drafter for the size you want to accelerate.

This is not the same as DeepSeek-V3-style MTP training objectives. Those change the training-time loss to predict multiple tokens at once, even if inference still emits one at a time. Co-designed drafters are an architectural artifact at inference time.

Speedups depend on the drafter agreeing with the target. Predictable patterns (code, structured output, repeated phrases) hit the high end of the curve. Highly creative or long-tail outputs land closer to 1×.

If you run Gemma 4 locally via Ollama or MLX, add the drafters.

References


About Sébastien

Ready to get to the next level?

Found this valuable? Share it with someone who needs it.

Join 6,000+ readers. Get practical systems for knowledge & AI. Free.

Subscribe ✨

Free: Knowledge System Checklist

A clear roadmap to building your own knowledge system. Subscribe and get it straight to your inbox.

6,000+ readers. No spam. Unsubscribe anytime.

Subscribe