DFlash 2 vs EAGLE-3 vs MTP¶
Canonical comparison of the three speculative-decoding draft paths you can actually deploy in 2026: the DFlash 2 block-diffusion drafter (DFlash 2 topic), the EAGLE-3 external autoregressive drafter, and native multi-token prediction (MTP) modules shipped with the model. Verification is identical across all three — the entire difference is in how tokens get drafted.
TL;DR¶
| DFlash 2 | EAGLE-3 | Native MTP | |
|---|---|---|---|
| What it is | External block-diffusion drafter + candidate path selector + two-tap dynamic convolutions | External autoregressive drafter conditioned on target features at its input | Multi-token prediction module shipped inside the target model |
| Drafting | Whole block (K=8-16) in one forward pass. Draft cost nearly flat in K | One token per forward pass. Draft cost grows linearly with K | One token per forward pass (for example, Qwen3.8: built-in 7-token serial draft) |
| Target conditioning | KV injection of target hidden states into every draft layer | Target features at the drafter input only (fades with depth) | Trained jointly with the model |
| Acceptance (best measured mean) | 5.97 (Qwen3.5-4B, 5-task mean) | 4.2-4.3 on comparable 4B-class rigs (5-layer drafter) | 4.54 (same Qwen3.5-4B rig). 4.28 (Qwen3.8-27B) |
| Speedup @ batch 1 | 2.67-3.43x (H200, Qwen3.8-27B) | 2.1-2.2x (LMSYS rig). 1.30x on TPU v5p out-of-box (K=2) | 1.96-2.59x (H200, Qwen3.8-27B) |
| Speedup @ concurrency 32 | 1.01-1.45x (SGLang model card). 2.20x on GSM8K in vLLM PR #52816 (different rig) | Not published on comparable rigs | 0.77-1.04x — can fall below baseline |
| Coverage | 4 standalone incoai DFlash 2 drafters (Qwen3.8-27B, Muse-Glimmer-30B, GLM-5.3, GLM-5.3-Flash), plus DFlash 2 drafts bundled in Inco Splash packages (Qwen3.8-27B, Qwen3.6-35B-A3B; Splash only). ~25 DFlash v1 drafters across Qwen/Gemma/Kimi/gpt-oss/Llama/MiniMax/GLM | Widest community-checkpoint ecosystem | Every model that ships an MTP head |
| Extra artifact to trust | Drafter checkpoint (new orgs, per-checkpoint licenses: Apache-2.0 for Qwen3.8-27B-DFlash2, CC BY-NC-ND 4.0 for GLM-5.3-Flash-DFlash2; read the card for the others) | Community drafter checkpoint | None — part of the model |
| Best for | Max speedup at interactive concurrency where a drafter exists | Broad model coverage with mature mainline engine support | Zero-effort baseline. Models without external drafters |
Layer Mapping¶
Forcing the three into explicit layers resolves most confusion: MTP and EAGLE-3 both draft serially and differ in who trained the drafter and what it sees. DFlash 2 changes the drafting execution model itself.
flowchart TB
subgraph DRAFTLAYER["Draft layer (the actual difference)"]
direction LR
MTP["Native MTP module<br/>in-model weights, serial K tokens<br/>trained with the target"]
EAGLE["EAGLE-3 drafter<br/>1-5 layers, serial<br/>target features at input"]
DF2["DFlash 2 drafter (2B, 5 layers)<br/>one pass, whole block<br/>KV injection + selector + conv"]
end
subgraph RUNTIME["Runtime layer"]
ENG["SGLang (Spec V2 overlap) / vLLM (Speculators) / llama.cpp<br/>TensorRT-LLM (DFlash 2 in 1.3.0rc28 pre-release)<br/>ollama, oMLX (branch/fork) / Splash (Inco, Apple Silicon)<br/>vLLM tpu-inference (DFlash v1 only)"]
end
subgraph VERIFY["Verification layer (identical for all three)"]
TV["Target LLM forward pass over draft block"]
RS["Rejection sampling -> target distribution, lossless"]
TV --> RS
end
MTP --> ENG
EAGLE -->|"target features at input"| ENG
DF2 -->|"target hidden states -> draft KV cache"| ENG
ENG --> TV
Architecture Comparison¶
Native MTP¶
The model's own extra prediction head(s), trained jointly with the target (Qwen3.8 ships a seven-token MTP path. Gemma 4 and DeepSeek-V4 ship MTP modules per LMSYS). Drafting is autoregressive — the head proposes tokens one at a time — so per-cycle draft cost grows with the speculation depth. It adds zero extra artifacts: the drafter is part of the model release, which makes it the only path with no supply-chain addition — but also the performance floor (see the TL;DR and Rig 1 for its acceptance and high-concurrency numbers).
EAGLE-3¶
The mature external-drafter design (arXiv:2503.01840): a small autoregressive transformer drafts sequentially, conditioned on the feature forecasts of the target model at the input of the drafter. That conditioning fades with drafter depth, so practical drafters are shallow (1 layer typical, 5-layer variants appear in ablations), and out-of-the-box speculation depths are small (K=2 on the TPU benchmark) because draft cost is linear in K. Strengths are ecosystem, not peak numbers: the widest set of community checkpoints, years of mainline engine support, and well-understood behavior (measured results in Rigs 3 and 4).
DFlash 2¶
External drafter, different execution model: a block-diffusion backbone predicts every position of the block in one non-causal forward pass, conditioned by injecting target hidden states through its own KV projection into every draft layer — the injection keeps deep drafters accurate, which input-only conditioning cannot. DFlash 2 adds a pairwise candidate path selector and two-tap dynamic convolutions to fix the two measured loss modes of parallel drafting (incoherent picks and suffix decay) at ~1.3% combined cycle latency. Component-level detail lives in DFlash 2 architecture. Draft cost is nearly flat in block size, which is why it exploits K-flat verification on datacenter hardware. Measured acceptance means: 5.97 (Qwen3.5-4B) and 4.80 (Qwen3.8-27B) — ahead of MTP and DSpark on every task, and +1.05 tokens over DFlash v1.
Measured Head-to-Heads (by rig)¶
Rig honesty matters here: no single published table contains all three methods on one rig. These are the four rigs with direct comparisons. Treat cross-rig rows as indicative, not equal.
Rig 1 — H200, Qwen3.8-27B, block 8 (Inco blog + HF model card. MTP and DFlash 2, no EAGLE-3):
| Concurrency | MTP | DFlash 2 |
|---|---|---|
| 1 | 1.96-2.59x | 2.67-3.43x |
| 8 | 1.74-2.19x | 2.27-2.85x |
| 32 | 0.77-1.04x | 1.01-1.45x |
Acceptance (mean of 5 tasks): MTP 4.28, DFlash 2 4.80.
Rig 2 — LMSYS Qwen3-4B-class, 5-layer drafters (EAGLE-3 vs DFlash v1):
| Task | EAGLE-3: acceptance / speedup | DFlash: acceptance / speedup |
|---|---|---|
| GSM8K | 4.2 / 2.1x | 4.2 / 3.3x |
| HumanEval | 4.3 / 2.2x | 4.0 / 3.2x |
| MT-Bench | 3.1 / 1.4x | 3.0 / 2.2x |
Same acceptance, ~1.5x the speedup — the win is drafting cost, not guess quality. With the selector+conv of DFlash 2 (Qwen3.5-4B rig), acceptance rises to 5.97 vs the 4.54 of MTP on the same task set.
Rig 3 — TPU v5p, Llama-3.1-8B (UCSD/Google. Out-of-box checkpoints, K=10 vs K=2): DFlash 2.29x vs EAGLE-3 1.30x end-to-end serving speedup.
Rig 4 — NVIDIA Blackwell, Speed-Bench, interactivity-matched (DFlash v1): avg interactivity speedup gpt-oss-120b: DFlash 2.3x vs EAGLE-3 1.7x. Llama-3.1-8B: 2.8x vs 2.2x. At 500-600 tok/s/user, DFlash delivers >15x AR throughput on 8x B300 — 1.5x higher than EAGLE-3 at the same point.
Caveat
Published EAGLE-3 comparisons are against DFlash v1, not DFlash 2. Because DFlash 2 strictly improves v1 (+1.05 tokens acceptance, same design), the v1-vs-EAGLE-3 gap is a floor for DFlash 2 — but no same-rig DFlash 2 vs EAGLE-3 table had been published as of 2026-09-27 (the DFlash 2 launch numbers compare against MTP, not EAGLE-3). Watch for third-party replications.
Which One Should I Pick?¶
The decision mostly turns on drafter availability, license, engine release status and traffic shape, not on peak benchmark numbers.
flowchart TD
Q1{"Official DFlash 2 or DFlash v1 drafter<br/>for your exact target checkpoint?"}
Q1 -->|No| Q4{"EAGLE-3 checkpoint available?"}
Q4 -->|Yes| EA["EAGLE-3"]
Q4 -->|No| Q5{"Model ships an MTP head?"}
Q5 -->|Yes| MT["Native MTP"]
Q5 -->|No| TR["Train a drafter<br/>(NeMo AutoModel) or run without speculation"]
Q1 -->|Yes| Q2{"Drafter license fits your use?<br/>(some are CC BY-NC-ND)"}
Q2 -->|No| Q4
Q2 -->|Yes| Q3{"Stable engine release with DFlash 2?<br/>SGLang 0.5.19+, vLLM 0.28.0+, llama.cpp, Splash"}
Q3 -->|"No: TensorRT-LLM stable, ollama, oMLX, TPU"| BR["DFlash 2 on a pre-release or branch build,<br/>DFlash v1 on TPU,<br/>or EAGLE-3/MTP on a stable release"]
Q3 -->|Yes| Q6{"Mostly batch 1-8 interactive traffic?"}
Q6 -->|Yes| DF["DFlash 2"]
Q6 -->|"No, concurrency 32+"| BM["Benchmark DFlash 2 vs MTP vs no speculation<br/>at real load before enabling"]
Decision Matrix¶
| Dimension | DFlash 2 | EAGLE-3 | Native MTP |
|---|---|---|---|
| Architecture | Block-diffusion, one pass, KV-injection conditioning | Autoregressive, input-only conditioning, shallow | In-model autoregressive head |
| Performance | Highest acceptance + speedup at batch 1-32 | Mid acceptance. Speedup capped by serial drafting | Mid acceptance. Degrades below baseline at high concurrency |
| Operations | Config swap. Best on SGLang Spec V2. Released in SGLang v0.5.19+, vLLM v0.28.0+ and llama.cpp (master since 2026-08-27). TensorRT-LLM only in pre-release 1.3.0rc28+. ollama/oMLX still branch/fork builds. Splash (vendor engine). TPU: DFlash v1 only |
Config swap. Mainline in the major engines | Zero config. Enabled by default path |
| Cost | +2B BF16 drafter resident. +1.3% cycle latency | Small shallow drafter | No extra artifact (weights ship with model) |
| Security | New drafter orgs. Check each checkpoint license. Lossless output | Mature checkpoint ecosystem. Lossless output | No new artifact. Lossless output |
| Ecosystem | 4 standalone incoai DFlash 2 drafters + 2 Splash-bundled drafts. ~25 v1 family. NeMo AutoModel training recipe | Widest checkpoint coverage. Years of engine support | Every MTP-shipping model |
| Lock-in Risk | Drafter per target. Official training code pending (Issue #1). Splash packages load only in Splash | Low — open algorithm, many checkpoints | None within a model. Changes per model family |
Weighted Decision Matrix¶
Scenario scored: you serve a model that has a published DFlash-family drafter. Interactive-heavy traffic (batch 1-8) on GPU. Scale 1-5, weights sum to 100%.
| Criterion | Weight | DFlash 2 | EAGLE-3 | MTP |
|---|---|---|---|---|
| Speedup at batch 1-8 | 25% | 5 | 3 | 3 |
| Behavior at concurrency 32 | 15% | 3 | 2 | 1 |
| Drafter availability (this scenario) | 15% | 5 | 4 | 3 |
| Engine support maturity | 15% | 4 | 5 | 5 |
| Rollout effort | 10% | 4 | 5 | 5 |
| Memory/overhead | 10% | 3 | 4 | 5 |
| Maturity / supply-chain risk | 10% | 3 | 4 | 5 |
| Weighted total (max 5.00) | 4.05 | 3.70 | 3.60 |
Weighted totals are the sum of weight x score (for DFlash 2: 1.25 + 0.45 + 0.75 + 0.60 + 0.40 + 0.30 + 0.30 = 4.05). The margin between DFlash 2 and EAGLE-3 is maturity vs peak: DFlash 2 wins every measured performance row but concedes points on engine maturity and ecosystem age. Flip the availability row (no DFlash drafter for your model: 5 -> 1) and DFlash 2 drops to 3.45, so EAGLE-3 wins at 3.70 — which is exactly the fallback rule in the Verdict.
Verdict¶
- Primary — DFlash 2 when a published drafter exists for your exact target model and traffic is interactive (batch 1-32): it leads acceptance and throughput on every published rig, is lossless, and migration is a config swap on SGLang/vLLM. Benchmark at your real concurrency before rollout — gains compress toward 1x by concurrency 32 (Rig 1).
- Fallback — EAGLE-3 when no DFlash-family drafter exists for your model (its checkpoint ecosystem is the widest), when the drafter license rules out your use (some DFlash 2 drafters are CC BY-NC-ND), or when you need an engine where DFlash 2 is not in a stable release (TensorRT-LLM stable, ollama, oMLX) or not ported (vLLM tpu-inference runs DFlash v1 only). Expect mid-2x speedups at low concurrency rather than ~3x.
- Baseline — native MTP as the zero-effort default and the A/B floor for any speculation rollout: it ships with the model, adds no artifact, but is the slowest of the three at batch 1 and measured below 1x at high concurrency (Rig 1) — do not assume it is free under load.
- Output quality is never the trade-off here: verification is the shared lossless layer (above), so an output audit valid under autoregressive decoding stays valid under all three.
Related Topics¶
- DFlash 2 topic — hub, explanation, and how-to guides
- LLM Fundamentals: speculative decoding — EAGLE, MTP, and draft-verify basics (EAGLE-3 and MTP have no dedicated topic folders yet)
- LLM Inference domain and comparisons index
Sources¶
URLs verified HTTP 200 on 2026-08-28. Engine and drafter status re-checked against the refreshed DFlash 2 reference on 2026-09-25.
- DFlash 2: Keep Drafting Parallel — Inco AI — DFlash 2 acceptance/throughput vs MTP and DSpark (H200 rig)
- incoai/Qwen3.8-27B-DFlash2 model card — concurrency 1/8/32 tables incl. MTP below 1x at 32
- DFlash: Block Diffusion for Flash Speculative Decoding — arXiv:2602.06036 — v1: >6x lossless, 2.5x over EAGLE-3
- EAGLE-3 — arXiv:2503.01840 — the autoregressive-drafter baseline
- Speculative decoding (original) — arXiv:2211.17192 — shared draft-and-verify foundation
- LMSYS: DFlash and Spec V2 — EAGLE-3 vs DFlash same-rig ablations. Qwen3.5-397B vs native MTP at all concurrencies
- NVIDIA: DFlash on Blackwell — EAGLE-3 vs DFlash interactivity tables
- Google: DFlash on TPUs — TPU v5p head-to-head, K=2 vs K=10
- z-lab/dflash — checkpoint catalog and engine integration surface
- TensorRT-LLM speculative decoding docs — MTP, EAGLE-3, and DFlash configs side by side
- incoai/GLM-5.3-Flash-DFlash2 — Hugging Face — post-launch DFlash 2 drafter and its CC BY-NC-ND 4.0 license
- vLLM PR #52816 — DFlash 2 in vLLM, 2.20x at concurrency 32 on GSM8K (author measurement)
- tensorrt-llm on PyPI — 1.3.0rc28 (2026-09-23) pre-release is the first build with DFlash 2. Stable 1.2.1 has no DFlash
- incoai/splash — GitHub — Inco's Apple Silicon engine with bundled DFlash 2 drafts (Apache-2.0)