← All writing

FIELD NOTES

Speculative decoding

Your best model is also your slowest. It crawls out one token at a time while the user waits.

You’ve already batched, cached, and moved to a real serving engine. Here’s the next lever: speculative decoding.

Hand-drawn infographic of speculative decoding: a big slow model, a small model guessing tokens, the big model checking them in one pass, wrong guesses overridden, 2–3× speedup

The idea is simple. A small, fast model guesses the next few tokens. Then the big model checks all the guesses in a single pass.

If the guesses were right — and on the easy stretches they often are — you just got several tokens for the price of one step.

If they were wrong, the big model overrides them. You lose no quality. The output draws from the exact same distribution as the big model alone. Under greedy decoding it’s token-for-token identical.

So you’re not trading accuracy for speed. You’re skipping the easy work.

Real systems see roughly 2 to 3× faster generation. One of the best free lunches in inference.