DFlash 2 How-to Guides¶
Context
Deployment recipes for every engine that runs DFlash 2 today (SGLang, vLLM, TensorRT-LLM, llama.cpp, ollama, oMLX, Splash), the dflash benchmark client, block-size tuning rules, drafter training via NVIDIA NeMo, and known operational pitfalls. Version matrices, flag tables, and benchmark numbers are in Reference.
Deployment¶
SGLang (recommended serving path, Spec V2 engine)¶
DFlash 2 support (PR #35371) first shipped in SGLang v0.5.19 (2026-09-04). The current release is 0.5.20 (2026-09-18). Install a tagged release instead of building from main. DFlash 2 uses the same DFLASH algorithm as v1: the engine picks the DFlash 2 model class from the drafter config:
pip install "sglang[all]>=0.5.19"
python -m sglang.launch_server \
--model-path Qwen/Qwen3.8-27B \
--speculative-algorithm DFLASH \
--speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2 \
--speculative-num-draft-tokens 8
For large MoE targets, the LMSYS/Modal/Z Lab launch command for Qwen3.5-397B-A17B (a DFlash v1 drafter, 8x B200) is the reference pattern:
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1
python -m sglang.launch_server \
--model-path Qwen/Qwen3.5-397B-A17B \
--trust-remote-code \
--speculative-algorithm DFLASH \
--speculative-draft-model-path modal-labs/Qwen3.5-397B-A17B-DFlash \
--speculative-dflash-block-size 8 \
--speculative-draft-attention-backend fa4 \
--attention-backend trtllm_mha \
--linear-attn-prefill-backend triton \
--linear-attn-decode-backend flashinfer \
--mamba-scheduler-strategy extra_buffer \
--tp-size 8 \
--max-running-requests 32 \
--cuda-graph-max-bs-decode 32 \
--cuda-graph-backend-prefill tc_piecewise \
--enable-flashinfer-allreduce-fusion \
--mem-fraction-static 0.8
The --linear-attn-* and --mamba-scheduler-strategy flags are there because Qwen3.5 is a hybrid linear-attention model. Drop them for a plain transformer target.
The model card for the Qwen3.5-397B drafter recommends block size 8 for higher-concurrency serving. Block size 16 gives longer acceptance and the best concurrency-1 throughput in most workloads. The same drafter is mirrored as lmsys/ and z-lab/Qwen3.5-397B-A17B-DFlash.
Migrating from EAGLE: change the speculative algorithm to DFLASH and point at a matching drafter — no application changes.
vLLM (via Speculators)¶
DFlash 2 support (PR #52816, merged 2026-08-21) is listed in the v0.28.0 release notes (2026-08-26). Any later release works (0.30.0 on 2026-09-22). num_speculative_tokens is the block size minus one:
pip install -U "vllm>=0.28.0"
vllm serve Qwen/Qwen3.8-27B \
--speculative-config '{
"method": "dflash",
"model": "incoai/Qwen3.8-27B-DFlash2",
"num_speculative_tokens": 7
}'
llama.cpp (local, single-request)¶
DFlash 2 support (PR #27342) merged to master on 2026-08-27. There is no new flag: DFlash 2 drafters are auto-detected from the GGUF and use the existing --spec-type draft-dflash. Build any post-merge master or release:
Reconvert older GGUFs
The PR notes that DFlash 2 GGUFs generated before 2026-08-27 must be reconverted to work with vision models. Download the drafter GGUF again if yours predates the merge.
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
# NVIDIA CUDA
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON
cmake --build build -j
# Apple Silicon
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON
cmake --build build -j
# Target and drafter GGUFs from local paths (the form used in PR #27342)
./build/bin/llama-server \
-m /models/Qwen3.8-27B-Q4_K_M.gguf \
-md /models/Qwen3.8-27B-DFlash2.gguf \
--spec-type draft-dflash \
--spec-draft-n-max 4
# Or pull the drafter from Hugging Face with -hfd <user>/<model>[:quant]
./build/bin/llama-server \
-m /models/Qwen3.8-27B-Q4_K_M.gguf \
-hfd incoai/Qwen3.8-27B-DFlash2-GGUF \
--spec-type draft-dflash \
--spec-draft-n-max 4
The PR benchmark (Qwen3.8-27B Q4_K_M, Apple M5 Pro) used --spec-draft-n-max 4 and measured 1.81x decode speed. Raise it toward 7 (block size 8) on GPUs where verification is cheap.
ollama (build from DFlash 2 branch)¶
Branch build
As of 2026-09-25, DFlash 2 in ollama lives only on the open PR #17865 ("mlx: add DFlash2 support"), which targets ollama's MLX engine on Apple Silicon. Mainline ollama releases have no DFlash 2 selector or convolution modules. Treat this as experimental. Reviewers reported that the default draft_num_predict of 0 silently disables speculation, and about 80% output agreement with non-speculative decoding. Set the draft depth explicitly and run the greedy diff test from the rollout checklist.
The build steps below follow the Inco AI launch instructions:
git clone https://github.com/ollama/ollama.git
cd ollama
git fetch origin pull/17865/head:dflash2
git switch dflash2
cmake -B build .
cmake --build build --parallel 8
TARGET="$(hf download mlx-community/Qwen3.8-27B-4bit)"
DRAFT="$(hf download incoai/Qwen3.8-27B-DFlash2)"
printf "FROM %s\nDRAFT %s\n" "$TARGET" "$DRAFT" > Modelfile
./ollama create qwen38-dflash2 --experimental --draft-quantize int4
./ollama serve
./ollama run qwen38-dflash2 --think high
oMLX (Apple Silicon GUI server)¶
- Install the prebuilt build: oMLX 0.6.2-zlab-dflash2 (arm64, signed).
- In the Model Downloader (
127.0.0.1:8891), downloadmlx-community/Qwen3.8-27B-4bitandincoai/Qwen3.8-27B-DFlash2. - Edit the target model in the Model Manager: enable DFlash, set draft model to
incoai/Qwen3.8-27B-DFlash2, enable draft quantization, runtime block size5, verify modedflash. - Load the target model and serve.
TensorRT-LLM¶
Supported on NVIDIA Blackwell and Hopper (NVIDIA published the gpt-oss-120b 8xB300 Pareto results with DFlash v1). The TensorRT-LLM speculative decoding docs document DFlashDecodingConfig. The same config loads DFlash 2 drafters with no extra arguments.
Pre-release only
As of 2026-09-25 the latest stable TensorRT-LLM (1.2.1) has no DFlash support. DFlashDecodingConfig exists in the 1.3.0 release candidates, and the DFlash 2 selector path is first tagged in 1.3.0rc28 (2026-09-23). Pin that pre-release (pip install "tensorrt-llm==1.3.0rc28") or a later one.
LLM API:
from tensorrt_llm import LLM
from tensorrt_llm.llmapi import DFlashDecodingConfig
speculative_config = DFlashDecodingConfig(
max_draft_len=7,
speculative_model="z-lab/Qwen3.8-27B-DFlash2",
attention_backend="TRTLLM", # SM100/SM103 only. Use "FA4" on Hopper (SM90) or the default "VANILLA"
)
llm = LLM("Qwen/Qwen3.8-27B", speculative_config=speculative_config)
trtllm-serve / trtllm-bench take the same options through --config config.yaml:
speculative_config:
decoding_type: DFlash
max_draft_len: 7
speculative_model: z-lab/Qwen3.8-27B-DFlash2
mask_token_id and target_layer_ids are read from the drafter config when unset. TensorRT-LLM warns when max_draft_len + 1 differs from the drafter's trained block_size. Keep them matched (7 for the Qwen3.8-27B drafter).
Splash (Inco AI, Apple Silicon)¶
Splash is Inco AI's own Apple-silicon engine (Apache-2.0, released 2026-09-18). Each model package bundles a 4-bit target and its DFlash 2 draft, so there is nothing to pair or tune. It needs an M3 or newer, macOS 26.4 or later, Homebrew, and 36 GB of unified memory (48 GB recommended):
brew install incoai/tap/splash
splash serve --model incoai/Qwen3.8-27B-Splash # or incoai/Qwen3.6-35B-A3B-Splash
The server binds 127.0.0.1:8000 and speaks OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages. Splash packages do not load in any other engine, and standard MLX or Transformers checkpoints do not load in Splash.
Commands & Recipes: the dflash Client¶
The repo's CLI is a lightweight client/benchmark harness — all speculation happens server-side. Requires Python 3.10+.
pip install dflash # client only (talks to a running server)
pip install "dflash[local]" # + local inference (MLX on Apple Silicon, Transformers on Linux)
Generate with the Transformers backend (DFlash 2 on Muse-Glimmer-30B):
dflash generate transformers \
--model meta-models/Muse-Glimmer-30B \
--draft z-lab/Muse-Glimmer-30B-DFlash2 \
--reasoning high --temperature 1 --top-p 0.95 --top-k 64 \
"How many positive whole-number divisors does 196 have?"
Generate locally on Apple Silicon (MLX, both models 4-bit — note block size 5, see tuning below):
dflash generate mlx \
--model mlx-community/Qwen3.8-27B-4bit \
--draft z-lab/Qwen3.8-27B-DFlash2 \
--draft-bits 4 --block-size 5 --reasoning xhigh \
"How many positive whole-number divisors does 196 have?"
Generate against any OpenAI-compatible SGLang/vLLM server:
dflash generate openai \
--base-url http://127.0.0.1:8000 --model Qwen/Qwen3.8-27B \
"How many positive whole-number divisors does 196 have?"
Benchmark (datasets: gsm8k, math500, humaneval, mbpp, mt-bench. Cached via HF Datasets):
dflash benchmark openai \
--base-url http://127.0.0.1:8000 --model Qwen/Qwen3.8-27B \
--dataset gsm8k --num-prompts 128 --concurrency 1 --reasoning xhigh \
--temperature 1 --top-p 0.95 --top-k 20
dflash benchmark mlx \
--model mlx-community/Qwen3.8-27B-4bit --draft z-lab/Qwen3.8-27B-DFlash2 \
--dataset gsm8k --max-samples 128 --reasoning xhigh --block-size 5 --draft-bits 4
Performance Tuning¶
| Knob | Guidance |
|---|---|
| Block size | 8 is the Qwen3.8-27B default (--speculative-num-draft-tokens 8 / num_speculative_tokens: 7 draft tokens). Muse Glimmer launch evals used 16. TPU measurements put the sweet spot at K=16 (>90% of theoretical max. 16→128 adds <1 token/step) |
| MLX quantized | Use block size <= 5 — the quantized matmul kernel of MLX loses efficiency at larger verify widths (both target and draft 4-bit) |
| Draft attention backend | --speculative-draft-attention-backend fa4 on SGLang for the non-causal block attention of the drafter |
| Spec engine | Prefer Spec V2 (overlap scheduler): >33% throughput gain over the V1 path at concurrency 32 (SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1 in current releases) |
| Reasoning effort | Qwen3.8: reasoning_effort low/medium/xhigh (default xhigh). Muse: reasoning_strength low/medium/high/xhigh (default high). The launch benchmarks ran at the highest setting (xhigh for Qwen3.8) |
| Concurrency expectations | Plan for ~3x at batch 1 compressing toward ~1x at concurrency 32 (full table: Reference — Benchmarks). Do not pay speculation overhead for pure-throughput batch workloads |
| Checkpoint pairing | The drafter must match the target exactly (incoai/Qwen3.8-27B-DFlash2 for Qwen/Qwen3.8-27B). A mismatched pair degrades or fails — drafters are per-target, not portable |
Training a Drafter for Your Own Model¶
No official training code ships in z-lab/dflash (Issue #1, opened 2026-01-06, is still open with no maintainer timeline). The documented path is the NeMo AutoModel recipe from NVIDIA (DFlash 2 draft model, trainer, and recipe added in NVIDIA-NeMo/Automodel PR #3605, merged 2026-08-22), which trains DFlash, DFlash 2, Domino, and JetSpec variants with the same scaffolding (frozen target, online hidden-state capture, DDP across GPUs):
torchrun --standalone --nproc_per_node=2 \
-m nemo_automodel.recipes.llm.train_dflash2 \
-c examples/speculative/dflash/qwen3_dflash2.yaml
Data rules that matter (from the recipe docs):
- Use chat-format prompts with responses regenerated by the target model — training is teacher-forced, so regenerating first prevents a train/inference distribution mismatch.
mask_token_idis required and must be a reserved, rarely-used token. Never reusepad(often aliased toeos) — it conflates the mask signal with content and quietly erodes acceptance. The inference runtime must fill block slots with the same id.- Objective: cross-entropy on block positions 1..K-1 weighted by position decay
w_k = exp(-(k-1)/loss_decay_gamma). Paper values track block size (gamma 7 for K=16, 5 for 10, 4 for 8). DFlash 2 adds a selector loss over top-k candidates weighted byselector_loss_weight.
DFlash 2-specific config fields: conv_kernel_size / conv_group_size (convolution taps and channels-per-correction), selector_rank / selector_top_k / selector_loss_weight (selector width, candidates per position, objective weight). Published drafters also diverge from their target via draft_sliding_window (1024) and a different attention shape (draft_num_attention_heads and friends). Supported DFlash 2 targets: Qwen3 dense and MoE, and the Qwen3.5 family including *ForConditionalGeneration variants (Qwen3.8-27B ships as a Qwen3.5-family model). Kimi K3 is registered for DFlash v1 training only. Watch the acceptance metrics (train/base_accept_len, train/candidate_recall, and the selector loss), not just the total loss. Full field list: Reference — NeMo AutoModel Training Config.
Rollout Checklist¶
Use this sequence to move a model from autoregressive decoding to DFlash 2 without surprises. The checks come from the monitoring and audit hooks and the concurrency data in Benchmarks.
flowchart TD
A{"Published DFlash 2 drafter<br/>for the exact target checkpoint?"} -->|No| B{"DFlash v1 drafter exists?"}
B -->|Yes| B1["Use DFlash v1<br/>same engine flags"]
B -->|No| B2["Use EAGLE-3 or native MTP<br/>or train with NeMo AutoModel"]
A -->|Yes| C["Check the model-card license<br/>(Apache-2.0 vs CC BY-NC-ND)"]
C --> D["Pin target and drafter HF revisions<br/>plus a tagged engine release"]
D --> E["Greedy diff test:<br/>AR output == DFlash output byte-for-byte"]
E -->|mismatch| F["Stop: engine or drafter bug"]
E -->|identical| G["Benchmark at real concurrency<br/>(1 / 8 / 32 / 64)"]
G -->|"speedup below ~1.1x at your load"| H["Leave speculation off<br/>for this traffic class"]
G -->|"clear gain"| I["Roll out, alert on<br/>acceptance-length drops"]
- Confirm a drafter exists for the exact target checkpoint (drafters are per-target) and read its license.
- Pin the target and drafter Hugging Face revisions and a tagged engine release.
- Run a fixed prompt set greedily with and without speculation. Outputs must be byte-identical.
- Benchmark with
dflash benchmark openai --concurrency <n>at the concurrency tiers you actually serve. - Record the artifact set (revisions, engine version) with the deployment and alert on mean acceptance-length drops.
Troubleshooting¶
- CUDA graph crash under load — z-lab/dflash Issue #146 (open) reports
cudaErrorIllegalAddresson vLLM 0.22.1 with DFlash plus CUDA graphs at 16+ concurrent requests. Reported workarounds:--enforce-eager(costs about 20-40% latency),--max-num-seqsat 32 or lower, or disabling DFlash. Re-test on a current vLLM release before relying on a workaround. - ollama branch shows no speedup — the PR #17865 build defaults
draft_num_predictto 0, which disables speculation. Set the draft depth explicitly. - llama.cpp drafter fails with a vision target — reconvert DFlash 2 GGUFs made before 2026-08-27.
- No speedup at high concurrency — expected behavior, not a bug. See Concurrency expectations in the tuning table.
- llama.cpp concurrency — speculative decoding in llama.cpp prevents concurrency scaling. Treat that path as single-user local inference only.
- Engine still on a PR ref — the launch-era instructions pinned vLLM
refs/pull/52816/head, SGLang PR #35371, and llama.cpp PR #27342. All three are merged now (vLLM v0.28.0+, SGLang v0.5.19+, llama.cppmasterafter 2026-08-27). Move to tagged releases. TensorRT-LLM needs 1.3.0rc28 or later (pre-release). Only ollama (PR #17865) and the oMLX fork remain branch builds. - Acceptance length regression after retraining — check the
mask_token_idaliasing rule and that responses were regenerated by the frozen target before training.
Cost Analysis¶
DFlash 2 decodes at close to 3x the speed of autoregressive decoding at roughly one third of the compute per token with identical output. Concretely (H200, Qwen3.8-27B, batch 1): ~69 tok/s autoregressive vs ~236 tok/s with DFlash 2 on GSM8K. The same GPU serves ~3.4x more interactive users at the same latency, or cuts per-token cost to roughly a third. The drafter itself adds ~2B parameters (BF16) of memory plus the ~1.3% cycle overhead.
Sources¶
- DFlash 2 blog — Inco AI — engine commands, oMLX steps
- z-lab/dflash README — CLI usage, checkpoint catalog, engine PR references
- NeMo AutoModel DFlash recipe — training commands, config fields, data rules
- LMSYS: DFlash and Spec V2 — SGLang flags, Spec V2 gains
- incoai/Qwen3.8-27B-DFlash2 model card — benchmark methodology
- oMLX 0.6.2-dflash2 release — Apple Silicon binary
- TensorRT-LLM speculative decoding docs —
DFlashDecodingConfig,decoding_type: DFlash - SGLang v0.5.19 release and vLLM v0.28.0 release — first releases listing DFlash 2 support
- llama.cpp PR #27342 — DFlash 2 auto-detection under
draft-dflash - NVIDIA-NeMo/Automodel PR #3605 — DFlash 2 draft model, trainer, and recipe
- modal-labs/Qwen3.5-397B-A17B-DFlash model card — block-size guidance for the MoE launch recipe
- ollama PR #17865 — DFlash 2 on the MLX engine, review notes
- z-lab/dflash Issue #146 — CUDA-graph crash details and workarounds
- incoai/splash — Splash install and model packages
- tensorrt-llm on PyPI — stable vs release-candidate versions