Skip to content

DFlash 2 vs EAGLE-3 vs MTP

Canonical comparison of the three speculative-decoding draft paths you can actually deploy in 2026: the DFlash 2 block-diffusion drafter (DFlash 2 topic), the EAGLE-3 external autoregressive drafter, and native multi-token prediction (MTP) modules shipped with the model. Verification is identical across all three — the entire difference is in how tokens get drafted.

TL;DR

DFlash 2 EAGLE-3 Native MTP
What it is External block-diffusion drafter + candidate path selector + two-tap dynamic convolutions External autoregressive drafter conditioned on target features at its input Multi-token prediction module shipped inside the target model
Drafting Whole block (K=8-16) in one forward pass. Draft cost nearly flat in K One token per forward pass. Draft cost grows linearly with K One token per forward pass (for example, Qwen3.8: built-in 7-token serial draft)
Target conditioning KV injection of target hidden states into every draft layer Target features at the drafter input only (fades with depth) Trained jointly with the model
Acceptance (best measured mean) 5.97 (Qwen3.5-4B, 5-task mean) 4.2-4.3 on comparable 4B-class rigs (5-layer drafter) 4.54 (same Qwen3.5-4B rig). 4.28 (Qwen3.8-27B)
Speedup @ batch 1 2.67-3.43x (H200, Qwen3.8-27B) 2.1-2.2x (LMSYS rig). 1.30x on TPU v5p out-of-box (K=2) 1.96-2.59x (H200, Qwen3.8-27B)
Speedup @ concurrency 32 1.01-1.45x (SGLang model card). 2.20x on GSM8K in vLLM PR #52816 (different rig) Not published on comparable rigs 0.77-1.04x — can fall below baseline
Coverage 4 standalone incoai DFlash 2 drafters (Qwen3.8-27B, Muse-Glimmer-30B, GLM-5.3, GLM-5.3-Flash), plus DFlash 2 drafts bundled in Inco Splash packages (Qwen3.8-27B, Qwen3.6-35B-A3B; Splash only). ~25 DFlash v1 drafters across Qwen/Gemma/Kimi/gpt-oss/Llama/MiniMax/GLM Widest community-checkpoint ecosystem Every model that ships an MTP head
Extra artifact to trust Drafter checkpoint (new orgs, per-checkpoint licenses: Apache-2.0 for Qwen3.8-27B-DFlash2, CC BY-NC-ND 4.0 for GLM-5.3-Flash-DFlash2; read the card for the others) Community drafter checkpoint None — part of the model
Best for Max speedup at interactive concurrency where a drafter exists Broad model coverage with mature mainline engine support Zero-effort baseline. Models without external drafters

Layer Mapping

Forcing the three into explicit layers resolves most confusion: MTP and EAGLE-3 both draft serially and differ in who trained the drafter and what it sees. DFlash 2 changes the drafting execution model itself.

flowchart TB
    subgraph DRAFTLAYER["Draft layer (the actual difference)"]
        direction LR
        MTP["Native MTP module<br/>in-model weights, serial K tokens<br/>trained with the target"]
        EAGLE["EAGLE-3 drafter<br/>1-5 layers, serial<br/>target features at input"]
        DF2["DFlash 2 drafter (2B, 5 layers)<br/>one pass, whole block<br/>KV injection + selector + conv"]
    end

    subgraph RUNTIME["Runtime layer"]
        ENG["SGLang (Spec V2 overlap) / vLLM (Speculators) / llama.cpp<br/>TensorRT-LLM (DFlash 2 in 1.3.0rc28 pre-release)<br/>ollama, oMLX (branch/fork) / Splash (Inco, Apple Silicon)<br/>vLLM tpu-inference (DFlash v1 only)"]
    end

    subgraph VERIFY["Verification layer (identical for all three)"]
        TV["Target LLM forward pass over draft block"]
        RS["Rejection sampling -> target distribution, lossless"]
        TV --> RS
    end

    MTP --> ENG
    EAGLE -->|"target features at input"| ENG
    DF2 -->|"target hidden states -> draft KV cache"| ENG
    ENG --> TV

Architecture Comparison

Native MTP

The model's own extra prediction head(s), trained jointly with the target (Qwen3.8 ships a seven-token MTP path. Gemma 4 and DeepSeek-V4 ship MTP modules per LMSYS). Drafting is autoregressive — the head proposes tokens one at a time — so per-cycle draft cost grows with the speculation depth. It adds zero extra artifacts: the drafter is part of the model release, which makes it the only path with no supply-chain addition — but also the performance floor (see the TL;DR and Rig 1 for its acceptance and high-concurrency numbers).

EAGLE-3

The mature external-drafter design (arXiv:2503.01840): a small autoregressive transformer drafts sequentially, conditioned on the feature forecasts of the target model at the input of the drafter. That conditioning fades with drafter depth, so practical drafters are shallow (1 layer typical, 5-layer variants appear in ablations), and out-of-the-box speculation depths are small (K=2 on the TPU benchmark) because draft cost is linear in K. Strengths are ecosystem, not peak numbers: the widest set of community checkpoints, years of mainline engine support, and well-understood behavior (measured results in Rigs 3 and 4).

DFlash 2

External drafter, different execution model: a block-diffusion backbone predicts every position of the block in one non-causal forward pass, conditioned by injecting target hidden states through its own KV projection into every draft layer — the injection keeps deep drafters accurate, which input-only conditioning cannot. DFlash 2 adds a pairwise candidate path selector and two-tap dynamic convolutions to fix the two measured loss modes of parallel drafting (incoherent picks and suffix decay) at ~1.3% combined cycle latency. Component-level detail lives in DFlash 2 architecture. Draft cost is nearly flat in block size, which is why it exploits K-flat verification on datacenter hardware. Measured acceptance means: 5.97 (Qwen3.5-4B) and 4.80 (Qwen3.8-27B) — ahead of MTP and DSpark on every task, and +1.05 tokens over DFlash v1.

Measured Head-to-Heads (by rig)

Rig honesty matters here: no single published table contains all three methods on one rig. These are the four rigs with direct comparisons. Treat cross-rig rows as indicative, not equal.

Rig 1 — H200, Qwen3.8-27B, block 8 (Inco blog + HF model card. MTP and DFlash 2, no EAGLE-3):

Concurrency MTP DFlash 2
1 1.96-2.59x 2.67-3.43x
8 1.74-2.19x 2.27-2.85x
32 0.77-1.04x 1.01-1.45x

Acceptance (mean of 5 tasks): MTP 4.28, DFlash 2 4.80.

Rig 2 — LMSYS Qwen3-4B-class, 5-layer drafters (EAGLE-3 vs DFlash v1):

Task EAGLE-3: acceptance / speedup DFlash: acceptance / speedup
GSM8K 4.2 / 2.1x 4.2 / 3.3x
HumanEval 4.3 / 2.2x 4.0 / 3.2x
MT-Bench 3.1 / 1.4x 3.0 / 2.2x

Same acceptance, ~1.5x the speedup — the win is drafting cost, not guess quality. With the selector+conv of DFlash 2 (Qwen3.5-4B rig), acceptance rises to 5.97 vs the 4.54 of MTP on the same task set.

Rig 3 — TPU v5p, Llama-3.1-8B (UCSD/Google. Out-of-box checkpoints, K=10 vs K=2): DFlash 2.29x vs EAGLE-3 1.30x end-to-end serving speedup.

Rig 4 — NVIDIA Blackwell, Speed-Bench, interactivity-matched (DFlash v1): avg interactivity speedup gpt-oss-120b: DFlash 2.3x vs EAGLE-3 1.7x. Llama-3.1-8B: 2.8x vs 2.2x. At 500-600 tok/s/user, DFlash delivers >15x AR throughput on 8x B300 — 1.5x higher than EAGLE-3 at the same point.

Caveat

Published EAGLE-3 comparisons are against DFlash v1, not DFlash 2. Because DFlash 2 strictly improves v1 (+1.05 tokens acceptance, same design), the v1-vs-EAGLE-3 gap is a floor for DFlash 2 — but no same-rig DFlash 2 vs EAGLE-3 table had been published as of 2026-09-27 (the DFlash 2 launch numbers compare against MTP, not EAGLE-3). Watch for third-party replications.

Which One Should I Pick?

The decision mostly turns on drafter availability, license, engine release status and traffic shape, not on peak benchmark numbers.

flowchart TD
    Q1{"Official DFlash 2 or DFlash v1 drafter<br/>for your exact target checkpoint?"}
    Q1 -->|No| Q4{"EAGLE-3 checkpoint available?"}
    Q4 -->|Yes| EA["EAGLE-3"]
    Q4 -->|No| Q5{"Model ships an MTP head?"}
    Q5 -->|Yes| MT["Native MTP"]
    Q5 -->|No| TR["Train a drafter<br/>(NeMo AutoModel) or run without speculation"]
    Q1 -->|Yes| Q2{"Drafter license fits your use?<br/>(some are CC BY-NC-ND)"}
    Q2 -->|No| Q4
    Q2 -->|Yes| Q3{"Stable engine release with DFlash 2?<br/>SGLang 0.5.19+, vLLM 0.28.0+, llama.cpp, Splash"}
    Q3 -->|"No: TensorRT-LLM stable, ollama, oMLX, TPU"| BR["DFlash 2 on a pre-release or branch build,<br/>DFlash v1 on TPU,<br/>or EAGLE-3/MTP on a stable release"]
    Q3 -->|Yes| Q6{"Mostly batch 1-8 interactive traffic?"}
    Q6 -->|Yes| DF["DFlash 2"]
    Q6 -->|"No, concurrency 32+"| BM["Benchmark DFlash 2 vs MTP vs no speculation<br/>at real load before enabling"]

Decision Matrix

Dimension DFlash 2 EAGLE-3 Native MTP
Architecture Block-diffusion, one pass, KV-injection conditioning Autoregressive, input-only conditioning, shallow In-model autoregressive head
Performance Highest acceptance + speedup at batch 1-32 Mid acceptance. Speedup capped by serial drafting Mid acceptance. Degrades below baseline at high concurrency
Operations Config swap. Best on SGLang Spec V2. Released in SGLang v0.5.19+, vLLM v0.28.0+ and llama.cpp (master since 2026-08-27). TensorRT-LLM only in pre-release 1.3.0rc28+. ollama/oMLX still branch/fork builds. Splash (vendor engine). TPU: DFlash v1 only Config swap. Mainline in the major engines Zero config. Enabled by default path
Cost +2B BF16 drafter resident. +1.3% cycle latency Small shallow drafter No extra artifact (weights ship with model)
Security New drafter orgs. Check each checkpoint license. Lossless output Mature checkpoint ecosystem. Lossless output No new artifact. Lossless output
Ecosystem 4 standalone incoai DFlash 2 drafters + 2 Splash-bundled drafts. ~25 v1 family. NeMo AutoModel training recipe Widest checkpoint coverage. Years of engine support Every MTP-shipping model
Lock-in Risk Drafter per target. Official training code pending (Issue #1). Splash packages load only in Splash Low — open algorithm, many checkpoints None within a model. Changes per model family

Weighted Decision Matrix

Scenario scored: you serve a model that has a published DFlash-family drafter. Interactive-heavy traffic (batch 1-8) on GPU. Scale 1-5, weights sum to 100%.

Criterion Weight DFlash 2 EAGLE-3 MTP
Speedup at batch 1-8 25% 5 3 3
Behavior at concurrency 32 15% 3 2 1
Drafter availability (this scenario) 15% 5 4 3
Engine support maturity 15% 4 5 5
Rollout effort 10% 4 5 5
Memory/overhead 10% 3 4 5
Maturity / supply-chain risk 10% 3 4 5
Weighted total (max 5.00) 4.05 3.70 3.60

Weighted totals are the sum of weight x score (for DFlash 2: 1.25 + 0.45 + 0.75 + 0.60 + 0.40 + 0.30 + 0.30 = 4.05). The margin between DFlash 2 and EAGLE-3 is maturity vs peak: DFlash 2 wins every measured performance row but concedes points on engine maturity and ecosystem age. Flip the availability row (no DFlash drafter for your model: 5 -> 1) and DFlash 2 drops to 3.45, so EAGLE-3 wins at 3.70 — which is exactly the fallback rule in the Verdict.

Verdict

  • Primary — DFlash 2 when a published drafter exists for your exact target model and traffic is interactive (batch 1-32): it leads acceptance and throughput on every published rig, is lossless, and migration is a config swap on SGLang/vLLM. Benchmark at your real concurrency before rollout — gains compress toward 1x by concurrency 32 (Rig 1).
  • Fallback — EAGLE-3 when no DFlash-family drafter exists for your model (its checkpoint ecosystem is the widest), when the drafter license rules out your use (some DFlash 2 drafters are CC BY-NC-ND), or when you need an engine where DFlash 2 is not in a stable release (TensorRT-LLM stable, ollama, oMLX) or not ported (vLLM tpu-inference runs DFlash v1 only). Expect mid-2x speedups at low concurrency rather than ~3x.
  • Baseline — native MTP as the zero-effort default and the A/B floor for any speculation rollout: it ships with the model, adds no artifact, but is the slowest of the three at batch 1 and measured below 1x at high concurrency (Rig 1) — do not assume it is free under load.
  • Output quality is never the trade-off here: verification is the shared lossless layer (above), so an output audit valid under autoregressive decoding stays valid under all three.

Sources

URLs verified HTTP 200 on 2026-08-28. Engine and drafter status re-checked against the refreshed DFlash 2 reference on 2026-09-25.