Learn/Core Concept How does speculative decoding speed inference? Speculative decoding uses a small model to draft tokens, then a larger model verifies them in batches. It can cut latency without changing output quality, useful when serving interactive APIs on shared GPUs. DistillationLatency |