Skip to content

LLM Inference

Domain summary

Techniques, engines, and serving infrastructure for making large language model generation faster and cheaper — speculative decoding, KV-cache engineering, scheduling, and hardware-aware serving.

This domain covers the optimization layer that sits between a trained model and its users. Since 2025 the field consolidated around a small set of levers: parallel (non-autoregressive) drafting, target-conditioned draft models, overlap scheduling that hides host latency, and block sizes tuned to hardware verification costs. Agent workloads — long, token-hungry, latency-sensitive — are the demand driver. See AI Agents for the consumer side.

Topics

Topic What It Is Latest Version License Status
DFlash 2 Block-diffusion speculative-decoding drafter (Inco AI, on Z Lab's DFlash) with a candidate path selector and two-tap dynamic convolutions. Lossless, 2.7-3.4x autoregressive throughput at batch 1 (H200, Qwen3.8-27B) dflash client 0.1.0 (2026-08-18). Engines: SGLang 0.5.19+, vLLM 0.28.0+ Code MIT. Drafter weights per checkpoint (Apache-2.0 to CC BY-NC-ND 4.0) active (announced Aug 2026)

Related foundations live in the AI Agents domain: LLM Fundamentals covers attention, the KV cache, quantization, and the speculative decoding overview this domain builds on.

How Speculative Serving Fits Together

The domain has one topic so far, so this map shows how it relates to its alternatives and runtimes. Every speculation path in this domain shares one loop: a drafter proposes a block, the serving engine batches and schedules it, and the frozen target verifies it in one forward pass. Only the drafting stage differs between DFlash 2, EAGLE-3, and native MTP.

flowchart LR
    REQ["Client request<br/>(OpenAI-compatible API)"] --> SCHED["Engine scheduler<br/>SGLang Spec V2 / vLLM / TensorRT-LLM /<br/>llama.cpp / Splash"]
    SCHED --> DRAFT{"Draft stage"}
    DRAFT -->|"one pass, whole block"| DF2["DFlash 2 drafter<br/>block diffusion + KV injection"]
    DRAFT -->|"serial, one token per pass"| EA3["EAGLE-3 drafter"]
    DRAFT -->|"serial, in-model head"| MTP["Native MTP module"]
    DF2 --> VER["Target LLM verify pass<br/>(all K positions at once)"]
    EA3 --> VER
    MTP --> VER
    VER --> RS["Rejection sampling<br/>(lossless)"]
    RS -->|"accepted prefix + bonus token"| SCHED
    RS --> OUT["Streamed tokens to client"]

Landscape Notes (2026)

  • Speculative decoding is the default. Every major engine (SGLang, vLLM, TensorRT-LLM, llama.cpp, ollama) ships at least one speculation path, and frontier open models (Qwen, Gemma, Kimi, DeepSeek) ship native MTP drafters. The differentiator is draft quality per unit of draft latency.
  • Parallel drafting displaced autoregressive drafting. The block-diffusion approach of DFlash — draft an entire block in one forward pass, condition on target hidden states via KV injection — set the pattern that successors (DFlash 2, Domino, DSpark, JetSpec) refine.
  • Verification is not the bottleneck on datacenter hardware. The TPU v5p measurements from UCSD ("K-flat verification") show verifying 1024 tokens costs nearly the same as verifying 16. Draft quality, not width, is the frontier.
  • Gains shrink with concurrency. Speculation speedups are largest at batch size 1-8 and compress toward 1x at concurrency 32+. Engine-level scheduling (overlap, Spec V2) matters as much as the drafter at high load.

Engine Landscape (2026-09)

Engine Niche Speculation Path (DFlash / DFlash 2)
SGLang Production serving. Spec V2 overlap scheduler is the current performance frontier --speculative-algorithm DFLASH + --speculative-draft-model-path. DFlash 2 merged in PR #35371 (2026-08-19), first shipped in v0.5.19 (2026-09-04)
vLLM Broadest model coverage. GPU via Speculators library, TPU via tpu-inference (JAX) --speculative-config JSON, "method": "dflash". DFlash 2 merged in PR #52816 (2026-08-21), listed in v0.28.0 release notes
TensorRT-LLM NVIDIA-tuned serving on Blackwell/Hopper DFlashDecodingConfig (LLM API) or decoding_type: DFlash in the --config YAML. The same config loads DFlash 2 drafters, but only in pre-release builds (first: 1.3.0rc28, 2026-09-23). Stable 1.2.1 has no DFlash
llama.cpp Local GGUF inference, single request --spec-type draft-dflash. DFlash 2 merged in PR #27342 (2026-08-27), auto-detected from the GGUF (no new flag)
ollama Easiest local UX DFlash 2 only on the PR #17865 branch (experimental draft models) as of 2026-09
oMLX Apple Silicon GUI server (MLX) DFlash via model-manager settings (Z Lab fork build 0.6.2-dflash2)
Splash (Inco AI) Apple Silicon engine (M3 or newer, macOS 26.4+), Apache-2.0 Launched 2026-09-18. Each Splash model package bundles a target and a DFlash 2 draft (Qwen3.8-27B, Qwen3.6-35B-A3B). Packages load only in Splash
vLLM TPU (tpu-inference) JAX serving on Google TPUs DFlash v1 only. No DFlash 2 port as of 2026-09-27 (checked in the 0.29.0 wheel)

Version status

Release-inclusion data comes from the SGLang and vLLM GitHub release notes and PyPI (checked 2026-09-25: vLLM 0.30.0, SGLang 0.5.20 and TensorRT-LLM 1.2.1 stable / 1.3.0rc28 pre-release are current). The full matrix is in DFlash 2 reference. Always confirm the release notes of the exact build you deploy.

Comparisons

Canonical comparison notes live in comparisons/ — currently DFlash 2 vs EAGLE-3 vs MTP: the three deployable speculative-decoding draft paths, measured head-to-heads, and a weighted decision matrix.

When to Use Which

A short version of the comparison's decision flowchart. Drafter availability and traffic shape decide more than peak benchmarks.

Situation Pick Why
A DFlash 2 drafter exists for your exact target, its license fits, traffic is mostly batch 1-8, engine is SGLang 0.5.19+, vLLM 0.28.0+ or llama.cpp DFlash 2 Highest measured acceptance and speedup (2.67-3.43x at concurrency 1 on the H200 rig)
No DFlash-family drafter for your model, or the drafter license (for example CC BY-NC-ND 4.0) rules out your use EAGLE-3 Widest community checkpoint coverage, mainline in every major engine
You need speculation with no extra artifact, or want an A/B floor Native MTP Ships inside the model. Can fall below 1x at concurrency 32
TensorRT-LLM stable, ollama, oMLX, or TPU serving EAGLE-3 or MTP (or DFlash v1 on TPU) DFlash 2 is pre-release (TensorRT-LLM 1.3.0rc28), branch/fork builds (ollama, oMLX) or not ported (tpu-inference)
Local single-user inference on Apple Silicon DFlash 2 via Splash, oMLX or llama.cpp Splash bundles target + draft. MLX quantized targets need block size <= 5
Heavy concurrency (32+) Benchmark first Gains compress toward 1x. Engine scheduling (SGLang Spec V2) matters as much as the drafter

FAQ

  • Is speculative decoding output really identical? Yes when implemented correctly — rejection sampling guarantees the target distribution. Engine bugs, not the algorithm, are the historical integrity risk. See DFlash 2 security.
  • Why not just add a bigger draft model? Draft cost, not draft size, is the old constraint — parallel (block-diffusion) drafting flattened it. This shifted the frontier to draft quality per cycle.
  • Do I need a new drafter per model? Yes — drafters are per-target artifacts. Check the Hugging Face collections before you assume support for a given checkpoint.
  • Does any of this help at concurrency 64+? The Inco/model-card data stops at 32 (1.01-1.45x). The SGLang v0.5.19 release notes report DFlash 2 about 24% ahead of DFlash v1 at concurrency 64 (a relative number, not a speedup over autoregressive decoding). At high concurrency, engine scheduling usually matters more than the drafter.
  • Where do drafters come from? Official HF orgs (incoai/z-lab, plus vendor orgs like nvidia, RedHatAI, modal-labs, lmsys) or self-training via the NVIDIA NeMo AutoModel recipe — see DFlash 2 how-to guides.

Reading Paths

  • "Why is my LLM serving slow?" — start with DFlash 2. Speculative decoding is the single largest lever for interactive (batch 1-8) workloads.
  • "Is it safe to roll out?" — read the concurrency table in DFlash 2 benchmarks first, then the rollout checklist.
  • "Can I trust speculative output?" — DFlash 2 security covers losslessness as an output-integrity guarantee and the drafter supply chain.

Key Concepts

Concept Meaning
Speculative decoding A small drafter proposes tokens. The target LLM verifies the block in one pass and accepts the longest valid prefix — lossless by rejection sampling
Acceptance length Mean committed tokens per draft-verify cycle. The single best health/quality metric for a speculation setup
Block size (K) Tokens drafted per cycle. Draft cost is nearly flat in K on modern hardware ("K-flat verification")
KV injection Conditioning technique: target hidden states feed the KV cache of the drafter at every layer (the approach of DFlash). This keeps deep drafters accurate
Suffix decay Accuracy loss toward the end of a drafted block. Fixed by local mixing (the convolutions of DFlash 2) rather than depth
Interactivity Per-user output token rate (tok/s/user) — the latency-side metric of the throughput-latency Pareto curve
Overlap scheduling Hiding host-side scheduler work behind GPU work (SGLang Spec V2). Worth >30% at high concurrency
Drafter The auxiliary model itself. Per-target, typically 1-2B params, distributed via Hugging Face collections

Sources

Open Questions

  • When does TensorRT-LLM ship DFlash 2 in a stable release (pre-release 1.3.0rc28 only as of 2026-09-25)?
  • Will the vLLM tpu-inference port gain the DFlash 2 selector and convolution modules?
  • No same-rig DFlash 2 vs EAGLE-3 benchmark exists yet. Published EAGLE-3 head-to-heads are all against DFlash v1.
  • Will official drafter training code ship (z-lab/dflash Issue #1, open), or does NeMo AutoModel stay the only documented path?
  • Candidate topics not yet covered: EAGLE-3 and native MTP (no dedicated folders), and the engines themselves (SGLang, vLLM, TensorRT-LLM).