topic: system-design
author: Crashtech Editorial
date: Oct 7, 2026 · read: 2 min
updated: October 8, 2026
---
Speculative Decoding: Faster Tokens Without Changing the Target Distribution
A draft model proposes tokens and a target verifies them. Exact speculative sampling keeps the target distribution; speedups depend on the workload.
On this page
Conceptual architecture; arrows show relationships, not measured latency, energy or safety guarantees.
A proposal is not an answer
A draft model generates a short continuation. The target evaluates the candidate positions with causal attention, allowing several distributions to be computed in one pass. Tokens are accepted in sequence until a rejection, or the proposal is exhausted.
The Leviathan paper reported two-to-three-times acceleration in its T5-XXL setup. The Chen paper reported two-to-two-and-a-half-times acceleration for its Chinchilla experiment. These are study-specific measurements, not a universal guarantee for any pair of models or GPU. [1] [2]
The correction step is essential
Let p be the target distribution and q the proposal distribution at the current position. For a token sampled from q, the acceptance probability is min(1, p(x)/q(x)). On rejection, the exact algorithm samples from the normalized positive part of p minus q.
acceptance(x) = min(1, p(x) / q(x))
replacement(x) proportional to max(0, p(x) - q(x))
Sampling a rejected replacement directly from p is not the same algorithm and does not establish the distribution-preservation result. When every proposed token is accepted, an additional target token can be sampled.
The guarantee concerns a probability distribution under the algorithm’s assumptions. It does not mean every run produces the same string, nor does it turn an inaccurate target model into a correct one. [1] [2]
Why verification can pay off
At low concurrency, dense-model decoding can be limited by moving weights through memory. Processing several candidate positions together can reuse weights more effectively than repeated single-token passes. Drafting and rejected work still have costs.
A 70B model stored at two bytes per weight has about 140 GB of weights before other memory requirements. Dividing that by one GPU’s bandwidth is not a valid single-GPU latency prediction if the model does not fit on that GPU. Sharding and communication must be included.
Measure the serving configuration
Compare inter-token latency, throughput, acceptance rate and peak memory at realistic prompt lengths and concurrency. vLLM documents implementation support and workload-dependent behavior; check the selected release rather than assuming every speculative method is interchangeable. [3]
Longer drafts are not always better. The useful proposal length balances the cost of producing and verifying candidates against the number actually accepted. Benchmark that balance before enabling it across a serving fleet.
Frequently asked questions
Does speculative decoding always double speed?
No. Speed depends on proposal quality, drafting cost, verification cost, batch size and implementation. Some workloads see little gain or a slowdown.
What happens when a draft token is rejected?
In exact speculative sampling, a replacement is drawn from a corrected residual distribution proportional to max(0, p minus q), not simply from the original target distribution.
/* Comments */
Comments are offline right now — we reconnect automatically, nothing is lost.