Speculative Decoding
A large model generates text correctly but slowly. A smaller model can quickly propose several following pieces. You want to use its work while keeping the large model in control of the final result.
Speculative Decoding produces a cheaper proposed continuation and then verifies it with the target model. In the classic algorithm, we accept the initial sequence of accepted tokens — pieces of text — until the first rejection, then compute a correction.
The large model can assess a known proposal at several positions in parallel. A correct sampling variant uses specific acceptance and correction probabilities to preserve the target distribution. Merely testing “does this sound sensible?” does not ensure that.
The speedup depends on the proposal's cost, the number of acceptances, and the hardware. If the small model often disagrees with the large one, extra work can cancel the gain. The demonstration below shows preservation of probabilities, rather than a guaranteed execution time.
Mechanism source: Leviathan et al., §2, Algorithm 1.
Mechanism and details
This accelerates Decode while preserving the target model's distribution under the algorithm's conditions. Checking whether a proposal “looks good” is not enough.
Why a correction is needed
Let be the target distribution and the draft distribution for the same prefix and a shared token set. Both include the chosen sampling rules, such as Temperature and Top-p Sampling. We accept a proposal with probability:
A sampled proposal has . After rejection, we sample from the distribution proportional to . This step fills exactly the missing probability mass. If , nothing is rejected, and no correction is needed. Leviathan et al., §2.3 and proof A.1.
An original example: for tokens A, B, C, let and . The unconditional mass of accepted proposals is . The remaining 0.4 must be split equally between B and C. Sampling the correction from ordinary would instead yield .
The experiment shows the exact mass balance for one position, without simulating the whole loop or execution time. Preserving the distribution does not guarantee identical text for the same seed: the algorithms consume randomness differently, and computations have limited precision. This is discussed by Chen et al., §4.2 and §6.1.
Speedup depends on the acceptance rate, draft cost, and hardware. An overly expensive draft or many rejections can erase the savings even when distributional correctness is preserved.
I use AI-generated content as part of my daily learning process.