Skip to content

DFlash 2 Explanation

Context

This page has two parts. Architecture (first half) explains how DFlash 2 works: the one-pass block-diffusion draft inherited from DFlash (Z Lab, ICML 2026), the KV-injection conditioning that keeps a 2B-parameter drafter accurate, and the two DFlash 2 additions — the candidate path selector and the two-tap dynamic convolution — each fixing one measured loss mode of parallel drafting. Security (second half) covers losslessness as an integrity guarantee, the drafter supply chain, and the threat model. Benchmark tables live in Reference. This page explains why the numbers look the way they do.

Speculative Decoding Context

A speculation cycle has two phases. A small drafter proposes a block of future tokens. The frozen target LLM verifies the whole block in one forward pass and accepts the longest valid prefix (speculative decoding fundamentals cover the mechanism). Rejection sampling makes the scheme lossless. Speedup reduces to acceptance length x (verification cost / (draft cost + verification cost)). The two questions: how many tokens survive each cycle, and how cheap the guessing was.

Autoregressive drafters (EAGLE-3, native MTP) need K sequential forward passes to propose K tokens, so their draft cost grows linearly with block size and they run shallow drafters. The TPU v5p measurements from UCSD added the hardware punchline, K-flat verification. On datacenter accelerators, verifying 1024 drafted tokens costs nearly the same as verifying 16. Weight loading dominates, not attention math. Draft quality, not width, is the binding constraint — which is exactly the constraint DFlash 2 attacks.

Component Breakdown

Component Role Key Facts
Target LLM (frozen) Verifier. Also the context expert For example, Qwen3.8-27B, Muse-Glimmer-30B. Shares embed_tokens and lm_head with the drafter
DFlash drafter backbone One-pass block denoiser ~2B params (Qwen3.8-27B drafter), 5 layers, sliding-window attention (1024-token window), own attention shape independent of target
Block attention mask Non-causal in-block attention Position 0 holds the real anchor token. Positions 1..K-1 start as MASK and are denoised in parallel
KV injection path Conditioning on target context Target hidden states pass through the KV projection of the drafter into the KV cache of every draft layer — not just the input embedding (the EAGLE-3 way)
Candidate path selector (DFlash 2) Coherence without autoregression — fixes incoherent independent picks (detail below) Top-16 candidates per position, low-rank bilinear pair scoring. +2.0M params, +0.6% cycle latency
Two-tap dynamic convolution (DFlash 2) Suffix-decay fix (detail below) Two taps around every attention and MLP sublayer, mixing each position with its predecessor. +16.5M params (+3%), +0.7% cycle latency
DFlashWorker / DFlashDraftModel / DFlash2DraftModel SGLang runtime Draft worker drives the scheduler and wraps the target worker for verification passes (PR #22077, then Spec V2 in PR #23000). DFlash2DraftModel adds the CandidateSelector and reuses the DFLASH worker (PR #35371, released in v0.5.19). The selector walk runs as a Triton kernel
Speculators library + qwen3_dflash2 model vLLM runtime Connects the drafter to target hidden states inside the vLLM path. Swap EAGLE-3 for DFlash by config only. DFlash 2 modules (DFlash2DraftModel) in PR #52816 (vLLM v0.28.0+). The PR routes DFlash 2 checkpoints to Model Runner V2
DFlashDecodingConfig TensorRT-LLM runtime Captures hidden states from selected target layers (target_layer_ids) as cross-attention context for the drafter. Separate cross-attention backend (VANILLA, TRTLLM on SM100/SM103, FA4 on SM90). DFlash 2 selector path first tagged in 1.3.0rc28 (pre-release)
TPU proposer (tpu-inference) JAX runtime Dual-cache: target keeps paged KV (Pallas kernels), drafter uses static on-device JAX arrays (PRs #1868-1870). DFlash v1 only as of 2026-09-25
llama.cpp draft-dflash C++ runtime A GGUF whose metadata sets selector_top_k > 0 switches to the DFlash 2 path. The selector reads its candidate lattice from the draft's hidden-state output instead of raw logits (PR #27342)

Draft-Verify Cycle

flowchart TB
    subgraph CTX["Context Capture (prefill / last verify)"]
        T0["Target LLM layers<br/>(frozen weights)"]
        HP["Target hidden states<br/>of context tokens"]
        T0 --> HP
    end

    subgraph DRAFT["Drafter — single forward pass"]
        KV["KV injection:<br/>draft KV projection of hidden states<br/>into every draft layer"]
        BB["5-layer block-diffusion backbone<br/>sliding window 1024, non-causal block mask"]
        CONV["Two-tap dynamic convolution<br/>before/after each attn + MLP sublayer"]
        HEAD["Shared frozen lm_head<br/>logits + top-16 candidates per position"]
        KV --> BB --> CONV --> HEAD
    end

    subgraph SEL["Path Selection (sequential, tiny)"]
        SC["Pairwise scorer<br/>U_t(b) + bilinear match A(a), B(b)<br/>gated by context H(h_t)"]
        WALK["Greedy walk or sampling<br/>from last verified token"]
        SC --> WALK
    end

    subgraph VER["Verification"]
        TV["Target LLM forward pass<br/>over draft block + next token"]
        RS["Rejection sampling<br/>(lossless, target distribution)"]
        TV --> RS
    end

    HP --> KV
    HEAD --> SC
    WALK -->|"draft block (K tokens)"| TV
    RS -->|"accepted prefix + bonus token"| CTX
    RS -->|"committed output"| OUT["Detokenizer / client"]

One Cycle, Step by Step

sequenceDiagram
    autonumber
    participant S as SGLang Scheduler
    participant DW as DFlashWorker
    participant DR as DFlash Drafter (2B)
    participant PS as Path Selector (+2M)
    participant T as Target LLM (frozen)

    S->>DW: forward_batch_generation(batch)
    DW->>DR: inject target hidden states (KV projection)
    DR-->>DW: whole block predicted in ONE pass<br/>top-16 candidates per position
    DW->>PS: candidate lists + last verified token
    PS-->>DW: one coherent path (K tokens)
    DW->>T: verify draft block in parallel
    T-->>DW: target logits for every position
    DW->>DW: rejection sampling (lossless)
    DW-->>S: accepted prefix + 1 bonus token

The counterintuitive SGLang detail (useful when reading traces): the draft worker is the one that talks to the scheduler — it wraps and calls the target worker when drafts are ready, not the other way around.

How It Works: The Two DFlash 2 Fixes

Problem 1 — Independent picks are incoherent (selection headroom)

Predicting every position independently means each pick is locally plausible but jointly incoherent — two neighbors can both pick "decoding", and the stutter dies at verification. The evidence that this is a selection problem, not a prediction problem, sits in the candidate lists of DFlash itself. The data below uses the 5-layer Qwen3-4B drafter on GSM8K, conditioned on all earlier positions being right:

Metric Pos 0 1 2 3 4 5 6 Acceptance length
Recall@1 85.4% 80.3% 79.4% 78.3% 77.5% 75.9% 72.9% 4.27
Recall@16 99.5% 97.3% 94.8% 92.6% 90.8% 89.4% 87.8% 6.79

An oracle picking from the top-16 reaches 6.79 vs 4.27 — 2.5 tokens of pure selection headroom. The path selector of DFlash 2 harvests part of it: keep the top-16 candidates per position and score every adjacent pair,

S_t(a, b) = U_t(b) + < A(a) ⊙ H(h_t), B(b) >

where U_t(b) is the logit of the drafter for candidate b, A and B are compact 256-dimensional token embeddings, and H(h_t) is a context gate deciding which parts of the predecessor-match count — in essence a low-rank bilinear attention over adjacent candidates. All pairs score in one parallel shot (no extra backbone or LM-head pass). Only the final walk over precomputed scores is sequential. Greedy follows the best successor. Sampling draws from the same scores. Rejection sampling restores the exact target distribution.

Results vs the DSpark-style sequential correction (5-layer Qwen3-4B, GSM8K. Overheads relative to plain DFlash):

Method Added params Latency Acceptance @ T=0 @ T=1
DFlash — — 4.27 3.78
+ DSpark correction +77.8M +9.6% 4.49 4.08
+ path selector (DFlash 2) +2.0M +0.6% 4.61 4.25

Roughly 40x fewer parameters and 16x lower latency overhead than the correction-head approach, with higher acceptance. Choosing is cheaper than predicting.

Problem 2 — Suffix decay is local (backbone, not selection)

Even the oracle decays down the block (99.5% → 87.8% across positions in the recall table): the candidates themselves run out. Depth fixes it indiscriminately. 3-, 5-, and 15-layer drafters are identical at position 0 and fan apart down the block. But ten extra layers cost +15.2% cycle latency and erase the efficiency edge of DFlash.

The attention of DFlash shows where the capacity is needed. It has two jobs — read the context before the block, and model dependencies inside it — but the attention share of the block falls from 30% in layer 1 to 8% in layer 5, with the remainder concentrating in a shrinking handful of heads. DFlash 2 splits the jobs: a dedicated module takes the within-block work. The design puts a two-tap dynamic depthwise convolution before and after each attention and feed-forward sublayer:

Conv_k(x)_t = k_t,0 ⊙ x_t + k_t,1 ⊙ x_{t-1}

Each coefficient combines a learned base kernel with a small correction computed from the current hidden state (one correction shared per 16 channels). The first position reads the representation of the last verified token. Every later position reads its predecessor. Information crosses the block while all positions still compute in parallel — and the module is block-local and stateless, so attention, the LM head, and verification are untouched.

Outcome: the convolutions buy what ten extra layers bought (overheads in the tables in this section) — five-layer DFlash + convolutions nearly matches 15-layer DFlash on suffix decay. Average within-block attention across layers 4-5 falls from 9.4% to 0.5%: the module has absorbed the local work, and attention returns to reading context. Suffix decay is mostly a local problem.

Combined System

Selector + convolutions together add 1.3% to the draft-verify cycle latency. Per-request mean acceptance length (lossless rejection sampling. Default sampling per model, block size 8 for Qwen3.8-27B, 16 for Muse Glimmer):

Target MTP DFlash DSpark DFlash 2
Qwen3.5-4B (mean of 5 tasks) 4.54 4.92 5.49 5.97
Qwen3.8-27B (vs native MTP) 4.28 — 3.62 4.80
Muse-Glimmer-30B (vs official DFlash) — 4.44 4.48 5.70

On Qwen3.5-4B that is +1.05 tokens per pass over DFlash (+21%) and +0.48 over DSpark. On both launch models DFlash 2 averages more than a full token ahead of DSpark. On MATH-500 the gain is visible position by position: DFlash 2 holds ~86% conditional acceptance to the last position while every baseline ends the block 6-9 points lower.

Internals Worth Knowing (Serving Path)

  • Immediate materialization of the draft KV projection. Target latents are projected by the drafter ahead of the draft forward pass rather than stored. This preserves KV space and radix-cache prefix sharing. SGLang implements this with a layer-batched linear projection plus a fused Triton kernel for norm+RoPE post-processing.
  • Overlap scheduling (SGLang Spec V2). Host-side cleanup of batch N-1 and KV allocation for batch N overlap with GPU work. This cuts host-device synchronization. Combining DFlash with Spec V2 improved throughput >33% (11.4 → 15.3 ktok/s, Qwen3-8B on B200 at concurrency 32).
  • TPU dual-cache. The non-causal block diffusion of DFlash is incompatible with paged attention, so the tpu-inference port runs the target on paged KV (Pallas kernels) and the drafter on static on-device JAX arrays. A metadata rework fixed "sequence length inflation", where draft state drifted from the accepted-token count of the target.

Performance Characteristics

The published numbers follow directly from the mechanics above. Full tables with test conditions are in Reference — Benchmarks.

  • Why batch 1 wins most. At low concurrency the target's forward pass is memory-bound: weights are loaded once per step whether it verifies 1 token or 8. Every extra accepted token is almost free, so DFlash 2 reaches 2.7-3.4x on H200 at concurrency 1 (Qwen3.8-27B).
  • Why gains compress under load. As concurrency rises, the verify pass becomes compute-bound and rejected draft tokens start to cost real FLOPs. The same rig drops to 1.01-1.45x at concurrency 32, and native MTP falls below 1x on several tasks. Engine and rig matter: vLLM PR #52816 reports 2.20x at concurrency 32 on GSM8K, so benchmark the stack you run.
  • Why width is not the lever. The TPU v5p study found that K=16 already captures >90% of the theoretical maximum and that raising per-position acceptance is 2-3x more valuable than growing K. DFlash 2 targets exactly that: the selector and convolutions raise acceptance per position without widening the block.
  • Why DFlash beats EAGLE-3 at equal acceptance. The LMSYS ablation isolates the two v1 techniques: diffusion drafting alone gives speed at lower acceptance, and KV injection alone gives acceptance at lower speed. Together they reach 3.2-3.3x vs 2.1-2.2x for a 5-layer EAGLE-3 drafter on the same Qwen3-4B rig.
  • Why vendor headlines differ. NVIDIA's "15x" is throughput at matched interactivity on an 8-GPU Pareto curve, not a batch-1 speedup. See Benchmark Caveats for the other source discrepancies.

Security

Context

DFlash 2 is a decoding-time optimization, not a network service — its security surface is different from that of an inference engine. The angles that matter: losslessness as an output-integrity property, the drafter-checkpoint supply chain, privacy characteristics of the draft/verify data flow, and the (small) set of ways the optimization itself can go wrong operationally.

Output Integrity: Losslessness as a Guarantee

The core security-relevant property of DFlash 2 is that it is provably output-preserving:

  • Greedy decoding produces token-for-token the same output as the target model alone.
  • Sampled decoding draws from the exact distribution of the target model via rejection sampling over draft proposals.
  • The selector and convolutions only change which candidates get proposed — verification and rejection sampling are untouched, so the committed distribution is unchanged regardless of drafter quality.

This matters for regulated or audit-sensitive deployments: adding DFlash 2 does not change model behavior, alignment characteristics, or output policy. An output audit run on autoregressive decoding remains valid under DFlash 2.

What can break the guarantee in practice is runtime correctness, not the algorithm:

  • z-lab/dflash Issue #146 (opened 2026-07-09, still open) reports cudaErrorIllegalAddress on vLLM 0.22.1 when DFlash runs with CUDA graphs at 16+ concurrent requests. The reporter suspects a buffer-sizing mismatch between DFlash buffers and CUDA-graph capture sizes. A crash is visible, but any engine bug that silently mis-schedules speculation (for example, wrong accepted-prefix accounting) can corrupt output.
  • Branch builds carry extra risk: a reviewer on ollama PR #17865 (DFlash 2 on MLX) reported about 80% output agreement with non-speculative decoding, attributed to numerical divergence in batched verification. Treat any engine that fails a greedy diff test as not lossless.
  • The "sequence length inflation" bug of the TPU port (draft state drifting from the accepted-token count of the target) is the canonical example of this bug class: fixed by synchronizing the proposer strictly with the true accepted token count.
  • Operational rule: diff-test after every engine or drafter upgrade (see Monitoring and Audit Hooks) — losslessness makes the check exact.

Supply Chain: Drafter Checkpoints and Build Provenance

DFlash 2 adds a second artifact to the model supply chain — the drafter — loaded alongside the target model into the serving process with code-execution potential (custom architecture code, trust-remote-code in SGLang configs).

Artifact Source Risk Posture Mitigation
Official drafters incoai/* and z-lab/* on Hugging Face (mirrored pairs, for example Qwen3.8-27B-DFlash2) Moderate — new org, fast-moving project. Licenses differ per checkpoint (Qwen3.8-27B-DFlash2 lists Apache-2.0. incoai/GLM-5.3-Flash-DFlash2 is CC BY-NC-ND 4.0, research and evaluation only) Pin exact revisions. Prefer the z-lab/ mirror when pairing incoai vs z-lab. Scan safetensors before load. Read the license on each model card
Community drafters for example RadixArk/Qwen3.8-27B-DSpark, DaoCloud/Muse-Glimmer-30B-DSpark, vendor drafter orgs (nvidia, RedHatAI, modal-labs, XiaomiMiMo, poolside, meta-models) Higher — third-party training provenance unknown Treat as untrusted model code. Review model cards for training data claims. Quarantine in a staging server first
GGUF drafts incoai/Qwen3.8-27B-DFlash2-GGUF (llama.cpp path) Moderate — conversion adds a transform step Verify conversion provenance. Compare greedy output against the safetensors path once
Engine builds Mainline since Aug-Sep 2026: vLLM v0.28.0+ (PR #52816), SGLang v0.5.19+ (PR #35371), llama.cpp (PR #27342, merged 2026-08-27). Pre-release only: TensorRT-LLM 1.3.0rc28. Still branch/fork builds: ollama PR #17865, oMLX signed dmg. Vendor engine: Inco Splash (Homebrew tap) Low for mainline releases. High for branch builds — code not through mainline review Use tagged releases where they exist. Pin exact commit SHAs for branch builds. Verify the oMLX signature
dflash pip package PyPI (dflash, MIT) Low — client/benchmark harness only Standard pip hash-pinning practice

Additional notes:

  • Drafter-target mismatch is a quality/availability issue, not an integrity one — a wrong drafter degrades speed and can crash the server, but lossless verification still filters any bad tokens. Do not rely on that as a safety net for code-level compromise, which verification cannot catch.
  • No known CVEs or published compromises of DFlash artifacts found as of 2026-09-25 (web search) — this reflects the age of the project (v1 Feb 2026), not audit depth. Re-check the issue tracker and advisories before production use.

Privacy and Data Flow

  • All computation stays in the serving process. Drafting, selection, and verification operate on the same prompts and KV state as ordinary decoding. No DFlash-specific data leaves the server. The dflash CLI documents no telemetry (its only outbound channel is a feedback form) — verify against the installed version if this matters to you.
  • No persistent cross-request state. Draft-side KV state is transient per cycle (Internals Worth Knowing). Nothing drafter-side persists beyond the request. Standard engine KV-cache isolation and prefix-sharing rules still govern cross-request leakage — DFlash does not weaken them by design, but engine-level cache-isolation bugs can now span both models.
  • The API perimeter is unchanged. Clients talk to the standard OpenAI-compatible endpoint. Apply the usual controls: API-key auth, TLS, network segmentation of the GPU serving VLAN, per-tenant rate limits.
  • Prompt content shapes acceptance, not exposure. Higher task predictability (math/code) raises acceptance. Nothing about the drafter makes prompt data more observable to other tenants.

Access Control and Hardening

  • Model-loading is the privileged operation. Whoever can point a SGLang/vLLM server at a drafter repo executes new model code with trust-remote-code. Restrict server launch and model-manager access (relevant for GUI servers like oMLX on shared machines — its admin dashboard binds to 127.0.0.1:8891 by default, keep it that way on multi-user hosts).
  • Resource controls bound speculation overhead. mem-fraction-static, max-running-requests, and cuda-graph-max-bs-decode cap the memory and scheduling footprint of the drafter. The drafter adds ~2B params (BF16) of resident memory.
  • Denial of budget, not denial of service. At high concurrency, speculation overhead compresses gains toward 1x — a misconfigured rollout can silently waste drafter memory and cycle time. Watch acceptance length (see Monitoring and Audit Hooks).

Threat Model Summary

Threat Vector Impact Likelihood Control
Malicious drafter checkpoint HF repo / mirror swap Code execution in server process Low today, grows with ecosystem Pin revisions, scan artifacts, staging server, prefer official mirrors
Tampered engine build Branch or PR-ref installs (ollama, oMLX fork) Arbitrary code execution Low Pin SHAs, move to mainline releases, verify signatures (oMLX dmg)
Silent output corruption Engine speculation bug (for example, #146-class, TPU seq-inflation-class, branch-build verify divergence) Integrity loss on committed tokens Low Greedy diff-testing vs autoregressive mode. Track the issue tracker. Acceptance-length monitoring
Cross-request data leakage Engine KV/radix-cache bug spanning target and draft KV Confidentiality breach Low (design preserves isolation) Standard engine isolation updates. Keep the engine patched
Budget waste (spec overhead) High-concurrency rollout without benchmarking Cost/throughput regression Medium Concurrency-tier benchmarking before rollout. Disable speculation for throughput-shaped traffic
Model unavailability Checkpoint-locked ecosystem. No training code for private fine-tunes Cannot accelerate custom models Certain today (Issue #1) NeMo training recipe. Z Lab/Modal engagement for custom drafters
Mask-token aliasing (training) Reusing pad/eos as mask_token_id in a self-trained drafter Quiet acceptance-length erosion (quality, not integrity) Medium Reserve a dedicated rarely-used token. Match the id at inference. Monitor train/accept_len

Monitoring and Audit Hooks

  • Track acceptance length per request — the natural health metric (mean committed tokens per verification step). Sudden drops indicate a wrong block size, a mismatched drafter, quantization drift, or a runtime bug. The dflash benchmark harness reports it directly.
  • Keep a lossless diff in the release checklist — run a fixed prompt set greedily through autoregressive and DFlash modes after every engine/drafter change. Outputs must be byte-identical, which makes output-integrity regression testing exact rather than statistical.
  • Record artifact provenance at deploy time — log the exact HF revision SHAs of target and drafter, the engine build (PR ref or release), and the oMLX/pkg versions alongside service version so any supply-chain incident maps to a precise artifact set.
  • Watch the upstream issue tracker — correctness-class reports (for example, CUDA-graph crashes) land as GitHub issues first. There is no separate security-advisory channel for the drafter ecosystem as of 2026-09-25.

Sources

Architecture:

Security: