FIELD NOTES
Speculative decoding
2026-07-15
Your best model is also your slowest. It crawls out one token at a time while the user waits.
You’ve already batched, cached, and moved to a real serving engine. Here’s the next lever: speculative decoding.

The idea is simple. A small, fast model guesses the next few tokens. Then the big model checks all the guesses in a single pass.
If the guesses were right — and on the easy stretches they often are — you just got several tokens for the price of one step.
If they were wrong, the big model overrides them. You lose no quality. The output draws from the exact same distribution as the big model alone. Under greedy decoding it’s token-for-token identical.
So you’re not trading accuracy for speed. You’re skipping the easy work.
Real systems see roughly 2 to 3× faster generation. One of the best free lunches in inference.