Skip to content

DFlash 2

Summary

DFlash 2 is the second-generation block-diffusion speculative-decoding drafter from Inco AI, built on the DFlash paper (Chen, Liang, Liu — Z Lab at UCSD, ICML 2026). It drafts an entire block of tokens in a single forward pass. Two cheap modules sit on top. A pairwise candidate path selector stitches independently predicted positions into one coherent path. A two-tap dynamic convolution stops draft accuracy from decaying toward the end of the block. Output is provably unchanged from the target model (lossless). Result: 16-25% more accepted tokens per verification pass for ~1.3% added cycle latency, and 2.7-3.4x autoregressive throughput at batch size 1 on H200 (Qwen3.8-27B). Mainline in SGLang 0.5.19+ and vLLM 0.28.0+ since Aug-Sep 2026.

Key Facts

Field Value
Full name DFlash 2: Keep Drafting Parallel (Inco AI, Aug 2026)
Lineage DFlash — Chen, Liang, Liu (Z Lab at UCSD), arXiv:2602.06036, ICML 2026
Type Speculative-decoding drafter (not a standalone LLM)
Latest Version (date) dflash client 0.1.0 (2026-08-18). Engines: SGLang 0.5.20 (2026-09-18), vLLM 0.30.0 (2026-09-22)
First engine releases with DFlash 2 vLLM v0.28.0 (2026-08-26), SGLang v0.5.19 (2026-09-04), TensorRT-LLM 1.3.0rc28 (2026-09-23, pre-release)
Code license MIT, (c) 2026 Z Lab (z-lab/dflash, dflash on PyPI)
Drafter licenses Per checkpoint: Apache-2.0 (Qwen3.8-27B-DFlash2), CC BY-NC-ND 4.0 (GLM-5.3-Flash-DFlash2)
Launch drafters Qwen3.8-27B (~2B params, BF16, block size 8) and Muse-Glimmer-30B (block size 16)
Later drafters incoai GLM-5.3 and GLM-5.3-Flash. Qwen3.8-27B and Qwen3.6-35B-A3B drafts bundled in Inco Splash packages
Headline numbers 2.7-3.4x AR throughput @ batch 1 (H200). +16-25% accepted tokens/pass. +1.3% cycle latency
Engines SGLang, vLLM, TensorRT-LLM, llama.cpp (mainline). ollama, oMLX (branch/fork). Splash (Inco, Apple Silicon)

Full version matrix, checkpoint catalog, and flag tables: Reference.

Architecture at a Glance

One DFlash 2 cycle: the drafter reads target hidden states through KV injection, drafts the whole block in one pass, the selector picks a coherent path, and the target verifies it losslessly. Full detail is in the Explanation.

flowchart LR
    TGT["Target LLM<br/>(for example Qwen3.8-27B)"] -->|"hidden states"| KVI["KV injection<br/>into every draft layer"]
    KVI --> BB["5-layer block-diffusion drafter<br/>+ two-tap dynamic convolutions"]
    BB -->|"top-16 candidates per position"| SEL["Candidate path selector<br/>(+2.0M params)"]
    SEL -->|"one K-token draft"| VER["Target verify pass<br/>+ rejection sampling"]
    VER -->|"accepted prefix + bonus token"| TGT

Evaluation

  • Why it is better: Autoregressive drafters (EAGLE-3, native MTP) pay a serial forward pass per draft token. This caps speculation depth. DFlash 2 keeps the one-pass parallel draft and recovers its two accuracy losses: incoherent top-1 picks (fixed by the path selector) and suffix decay (fixed by the convolution). In the Inco AI benchmarks it beats its own v1 drafter, native MTP, and community DSpark drafters on every task tested (GSM8K, MATH-500, HumanEval, MBPP, MT-Bench).
  • When it fits: Latency-sensitive, low-concurrency serving (batch size 1-8) — interactive coding, reasoning, and agent workloads. Local single-user inference on Apple Silicon (Splash, oMLX, llama.cpp, ollama branch).
  • When it does not fit: High-concurrency throughput serving, targets without a published drafter, and TPU serving (the vLLM tpu-inference port is DFlash v1 only as of 2026-09-25). The DFlash v1 collection lists ~25 target checkpoints across Qwen, Gemma 4, MiniMax, Kimi, gpt-oss, Llama, and GLM (z-lab README), far more than the DFlash 2 set. Native MTP even dips below 1x on some tasks at concurrency 32.
Pros Cons
Lossless — greedy output matches target exactly. Sampling preserves its distribution Needs a drafter trained per target model. No official training code in the repo ("checkpoint-locked", z-lab/dflash Issue #1). NeMo AutoModel is the documented training path
One-pass parallel drafting: draft cost is nearly independent of block size (K-flat verification on TPU v5p) Gains compress at high concurrency (1.01-1.45x at concurrency 32 on the SGLang/H200 rig)
Tiny overhead: +1.3% draft-verify cycle latency combined (per-module split in Explanation) Few DFlash 2 drafters: two at launch (Qwen3.8-27B, Muse-Glimmer-30B), plus incoai GLM-5.3 and GLM-5.3-Flash since
Mainline in SGLang v0.5.19+, vLLM v0.28.0+, and llama.cpp. TensorRT-LLM in 1.3.0rc28+ pre-releases ollama and oMLX are still branch/fork builds. llama.cpp speculation is single-request oriented
Beats DSpark correction with ~40x fewer added parameters and 16x lower latency overhead MLX quantized targets must drop to block size <= 5 (quantized matmul kernel efficiency)
  • Common Use Cases:
    • Agent serving stacks that need per-user interactivity (500-600 tok/s) at high concurrency. NVIDIA measured 15x higher throughput than autoregressive decoding at the same interactivity for gpt-oss-120b on 8x Blackwell Ultra (DFlash v1, TensorRT-LLM).
    • Local agentic coding on Apple Silicon: Qwen3.8-27B 4-bit + DFlash 2 drafter via Splash, oMLX, or the dflash MLX backend.
    • Production endpoints: CoreWeave's Kimi K2.7 Code endpoint (fastest for that model on Artificial Analysis at publication) runs a custom NVFP4 quantization plus a DFlash v1 drafter on vLLM by default. DFlash 2 is the upgrade path once a drafter for that target exists.
  • Licensing & Commercial Use: The z-lab/dflash code and the dflash PyPI package are MIT (c) 2026 Z Lab. Drafter checkpoints carry per-checkpoint licenses that differ: incoai/Qwen3.8-27B-DFlash2 lists Apache-2.0, while incoai/GLM-5.3-Flash-DFlash2 is CC BY-NC-ND 4.0 for research and evaluation, with commercial licensing through Inco. The Splash engine is Apache-2.0. Read each model card before commercial use.
  • Ecosystem & Data Connections: Runs inside SGLang (default Spec V2 engine), vLLM, TensorRT-LLM, llama.cpp, ollama (MLX engine branch), oMLX, and Inco's Splash. DFlash v1 also runs on vLLM TPU / tpu-inference (JAX). Engine versions and PRs are in Reference. Drafters are distributed through the Hugging Face collections incoai/dflash-2, z-lab/dflash, and z-lab/dflash-2.
  • Compatibility & Requirements: A matching DFlash 2 drafter for the exact target model. CUDA (Blackwell and Hopper), Apple Silicon (MLX, Splash on M3 or newer with macOS 26.4+). The drafter is not a standalone LLM — it requires a speculation-aware server. Python 3.10+ for the dflash CLI/benchmark client.
  • Latest Versions: DFlash paper arXiv:2602.06036 (v1 2026-02-05, v2 2026-05-28, ICML 2026). DFlash 2 announced Aug 2026 with two drafters. dflash client 0.1.0 (2026-08-18). SGLang 0.5.20 (2026-09-18) and vLLM 0.30.0 (2026-09-22) are the current engine releases. Inco Splash launched 2026-09-18. More than 3.5M Hugging Face downloads across DFlash models (Inco blog, Aug 2026).
  • Alternatives: EAGLE-3 (autoregressive drafter, input-only conditioning), native MTP modules (ships with the model, 7-token serial draft on Qwen3.8), DSpark (sequential correction heads, +77.8M params/+9.6% latency), Domino (serial GRU correction), JetSpec (causal in-block attention + forward-KL distillation), SpecDiff-2 (full diffusion LLM as drafter — high memory footprint), Medusa (parallel heads, older generation).
  • Migration & Lock-in Risks: Migration from EAGLE/MTP is a config change in SGLang/vLLM/TensorRT-LLM (swap algorithm + drafter path — no application refactoring). Lock-in is model-shaped, not vendor-shaped: each target model needs its own drafter, and third-party drafter training currently means NVIDIA NeMo AutoModel recipes, community SpecForge work, or Z Lab/Modal engagement. Splash packages are the exception: they load only in Splash. The algorithm itself is open (paper + MIT code).
  • Community Health: Active and industry-backed. Z Lab (UCSD) maintains the repo. SGLang/Modal/Z Lab co-engineered the serving path. The DFlash 2 engine ports for SGLang, vLLM, and llama.cpp came from the same author (SubSir) within eight days. Per the NVIDIA blog, the researchers released 20 DFlash checkpoints with recipes for Blackwell and Hopper GPUs. Google Cloud co-published the TPU port. Drafters were published by NVIDIA, Red Hat, Modal, Meta, Poolside, and Xiaomi. Known friction: training code not released (Issue #1, open since 2026-01-06), a CUDA-graph crash under load on vLLM 0.22.1 (Issue #146, open).

Topic Map

  • How-to Guides — serve with SGLang/vLLM/TensorRT-LLM/llama.cpp/ollama/oMLX/Splash, dflash CLI, block-size tuning, rollout checklist, training drafters with NeMo
  • Reference — release timeline, engine support matrix, drafter checkpoints and licenses, config keys and flags, benchmark tables and caveats
  • Explanation — draft/verify cycle, KV injection, path selector, two-tap convolution, why the benchmarks look the way they do, and the security model
  • Comparisons — index pointing to the canonical DFlash 2 vs EAGLE-3 vs MTP note

Sources

URLs verified HTTP 200 on 2026-08-28. Re-checked on 2026-09-25 by web search, GitHub page fetches, PyPI JSON, and upstream source files at release tags (several hosts block direct fetches from the research environment).

Questions

  • Will Z Lab/Inco release official drafter training code and end the "checkpoint-locked" limitation? Issue #1 has no maintainer timeline as of 2026-09-25. (How-to Guides documents the interim NeMo path.)
  • Partly answered: license terms are per checkpoint (Apache-2.0 for Qwen3.8-27B-DFlash2, CC BY-NC-ND 4.0 for GLM-5.3-Flash-DFlash2). Open: the terms for Muse-Glimmer-30B-DFlash2, GLM-5.3-DFlash2, and the z-lab mirrors — check each card.
  • Partly answered: SGLang v0.5.19 notes report ~24% over DFlash v1 at concurrency 64, and vLLM PR #52816 reports 2.20x over autoregressive at concurrency 32. Open: why the SGLang and vLLM concurrency-32 numbers differ so much, and the gain under chunked-prefill production traffic mixes.
  • When does TensorRT-LLM ship DFlash 2 in a stable release (currently 1.3.0rc28 pre-release only)?
  • Will the vLLM tpu-inference port gain the DFlash 2 selector and convolution modules?
  • Does a DFlash 2 drafter arrive for Qwen3.5-397B-A17B and Kimi-class MoE models, where v1 DFlash already beat native MTP at every concurrency tested?