Multi-Token Prediction and Its Descendants
Last Updated on August 25, 2026 by Editorial Team
Author(s): Enzo Lombardi
Originally published on Towards AI.
How MTP, DFlash, DFlash 2, and DSpark each attack the same bottleneck
Text generation is embarrassingly serial. To produce token number 200 a transformer needs token 199, which needs 198, and so on back to the prompt. Each of those steps drags the entire weight matrix out of memory to compute a single token. On modern accelerators that is a catastrophically bad deal: the arithmetic units sit idle while the memory bus does all the work. A single-user decode loop typically runs at a few percent of the hardware’s theoretical throughput, and no amount of extra FLOPs fixes it, because FLOPs were never the constraint.

After introducing speculative decoding as the core idea behind all the methods, the article explains the “shape of the trick” via an acceptance accounting model where verification stops at the first mismatch, making token acceptance depend on prefix products rather than independent per-position success. It then describes MTP, which adds multiple prediction heads to the same model and gains speed “for free” from a shared forward pass but suffers from mutually inconsistent lookahead that collapses accuracy with distance. DFlash instead drafts an entire block using a small block-diffusion-style draft model, achieving near-constant drafting cost per block but encountering “suffix decay” where acceptance degrades toward the block tail. DSpark keeps DFlash’s parallel backbone yet adds minimal autoregressive structure (and a confidence head) to adjust distributions and choose how many tokens to verify per request, including load-aware scheduling; it reports substantial improvements in accepted length and effective latency. DFlash 2 rejects DSpark’s sequential correction as too expensive, and instead fixes the tail in-parallel using lightweight convolutions and a reranking/path-selection mechanism, aiming for better acceptance gains with minimal latency and parameter overhead. Finally, the article ties the family tree to a single overarching lesson: speculative decoding trades arithmetic throughput for reduced wall-clock latency, and the best strategy is usually improving early guess quality and stopping at the right time rather than speculating further, with results that depend strongly on actual hardware and concurrency.
Read the full blog for free on Medium.
Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.
Published via Towards AI
Towards AI Academy
We Build Enterprise-Grade AI. We'll Teach You to Master It Too.
15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.
Start free — no commitment:
→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day
→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages
Our courses:
→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.
→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.
→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.
Note: Article content contains the views of the contributing authors and not Towards AI.