Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Read by thought-leaders and decision-makers around the world. Phone Number: +1-650-246-9381 Email: pub@towardsai.net
228 Park Avenue South New York, NY 10003 United States
Website: Publisher: https://towardsai.net/#publisher Diversity Policy: https://towardsai.net/about Ethics Policy: https://towardsai.net/about Masthead: https://towardsai.net/about
Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Founders: Roberto Iriondo, , Job Title: Co-founder and Advisor Works for: Towards AI, Inc. Follow Roberto: X, LinkedIn, GitHub, Google Scholar, Towards AI Profile, Medium, ML@CMU, FreeCodeCamp, Crunchbase, Bloomberg, Roberto Iriondo, Generative AI Lab, Generative AI Lab VeloxTrend Ultrarix Capital Partners Denis Piffaretti, Job Title: Co-founder Works for: Towards AI, Inc. Louie Peters, Job Title: Co-founder Works for: Towards AI, Inc. Louis-François Bouchard, Job Title: Co-founder Works for: Towards AI, Inc. Cover:
Towards AI Cover
Logo:
Towards AI Logo
Areas Served: Worldwide Alternate Name: Towards AI, Inc. Alternate Name: Towards AI Co. Alternate Name: towards ai Alternate Name: towardsai Alternate Name: towards.ai Alternate Name: tai Alternate Name: toward ai Alternate Name: toward.ai Alternate Name: Towards AI, Inc. Alternate Name: towardsai.net Alternate Name: pub.towardsai.net
5 stars – based on 497 reviews

Frequently Used, Contextual References

TODO: Remember to copy unique IDs whenever it needs used. i.e., URL: 304b2e42315e

Resources

Free: 6-day Agentic AI Engineering Email Guide.
Learnings from Towards AI's hands-on work with real clients.
Multi-Token Prediction and Its Descendants
Latest   Machine Learning

Multi-Token Prediction and Its Descendants

Last Updated on August 25, 2026 by Editorial Team

Author(s): Enzo Lombardi

Originally published on Towards AI.

How MTP, DFlash, DFlash 2, and DSpark each attack the same bottleneck

Text generation is embarrassingly serial. To produce token number 200 a transformer needs token 199, which needs 198, and so on back to the prompt. Each of those steps drags the entire weight matrix out of memory to compute a single token. On modern accelerators that is a catastrophically bad deal: the arithmetic units sit idle while the memory bus does all the work. A single-user decode loop typically runs at a few percent of the hardware’s theoretical throughput, and no amount of extra FLOPs fixes it, because FLOPs were never the constraint.

Multi-Token Prediction and Its Descendants

After introducing speculative decoding as the core idea behind all the methods, the article explains the “shape of the trick” via an acceptance accounting model where verification stops at the first mismatch, making token acceptance depend on prefix products rather than independent per-position success. It then describes MTP, which adds multiple prediction heads to the same model and gains speed “for free” from a shared forward pass but suffers from mutually inconsistent lookahead that collapses accuracy with distance. DFlash instead drafts an entire block using a small block-diffusion-style draft model, achieving near-constant drafting cost per block but encountering “suffix decay” where acceptance degrades toward the block tail. DSpark keeps DFlash’s parallel backbone yet adds minimal autoregressive structure (and a confidence head) to adjust distributions and choose how many tokens to verify per request, including load-aware scheduling; it reports substantial improvements in accepted length and effective latency. DFlash 2 rejects DSpark’s sequential correction as too expensive, and instead fixes the tail in-parallel using lightweight convolutions and a reranking/path-selection mechanism, aiming for better acceptance gains with minimal latency and parameter overhead. Finally, the article ties the family tree to a single overarching lesson: speculative decoding trades arithmetic throughput for reduced wall-clock latency, and the best strategy is usually improving early guess quality and stopping at the right time rather than speculating further, with results that depend strongly on actual hardware and concurrency.

Read the full blog for free on Medium.

Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.

Published via Towards AI


Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

6-Day Agentic AI Engineering Email Guide — one practical lesson per day

Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.