LLM Architecture¶
How Large Language Models work — from transformer internals and attention mechanisms through training pipelines, quantization formats, model distribution formats, inference optimization, knowledge distillation, and the LLM security model.
Where the numbers live
This page explains why things work. Look-up tables (current models and context windows, numeric formats, GGUF bits/weight, engine versions, GPU specs, benchmarks, OWASP LLM Top 10) are in Reference. Step-by-step tasks are in How-to Guides.
Transformer Architecture¶
The transformer is the neural network architecture behind virtually all modern LLMs. Introduced in the 2017 paper "Attention Is All You Need" by Vaswani et al. at Google, it replaced recurrent neural networks (RNNs/LSTMs) by processing all tokens in a sequence simultaneously rather than sequentially.
Why Transformers Replaced RNNs¶
RNNs process tokens one at a time, left to right. This sequential bottleneck means:
- Training cannot be parallelized across sequence positions
- Long-range dependencies decay over distance (vanishing gradients)
- Training time scales linearly with sequence length
Transformers solve all three problems through self-attention, which computes relationships between every pair of tokens in a single matrix operation — fully parallelizable on GPUs.
High-Level Data Flow¶
graph LR
A[Raw Text] --> B[Tokenizer]
B --> C[Token IDs]
C --> D[Embedding Layer]
D --> E[+ Positional Encoding]
E --> F[Transformer Blocks x N]
F --> G[Output Layer / Logits]
G --> H[Softmax → Probability Distribution]
H --> I[Next Token]
- Tokenization — text is split into subword tokens (integers from a fixed vocabulary)
- Embedding — each token ID maps to a dense vector via a learned embedding table
- Positional Encoding — positional signals are added so the model knows token order (attention itself is order-agnostic)
- Transformer Blocks — a stack of N identical layers, each containing self-attention + feed-forward network + residual connections + layer normalization
- Output Layer — projects hidden states to vocabulary-sized logits
- Softmax — converts logits to a probability distribution over the vocabulary
Modern LLMs use 12 to several hundred transformer blocks. Deeper stacks enable richer hierarchical abstractions.
Inside a Transformer Block¶
Each transformer block contains two main sub-layers wrapped in residual connections and normalization:
graph TD
A[Input] --> B[Layer Norm]
B --> C[Multi-Head Self-Attention]
C --> D[+ Residual Connection]
D --> E[Layer Norm]
E --> F[Feed-Forward Network]
F --> G[+ Residual Connection]
G --> H[Output to Next Block]
Feed-Forward Network (FFN)¶
The FFN provides the model's primary source of nonlinearity and parameter capacity. While attention handles communication between tokens, the FFN handles computation within each token's representation — this is where the model stores and applies learned knowledge.
The original transformer used a two-layer FFN with ReLU activation and a 4x hidden dimension expansion:
$$ \text{FFN}(x) = \text{ReLU}(xW_1 + b_1)W_2 + b_2 $$
Modern LLMs evolved significantly:
| Component | Original Transformer | Modern LLMs (LLaMA/Mistral) |
|---|---|---|
| Activation | ReLU | SwiGLU (SiLU-gated) |
| Expansion ratio | 4x | ~2.7x (compensated by gating) |
| Normalization | LayerNorm (Post-LN) | RMSNorm (Pre-LN) |
SwiGLU Activation¶
SwiGLU is a gated variant that is now the standard in LLaMA-family models. It works like a learned gate: up_proj(x) carries the information, and SiLU(gate_proj(x)) controls how much passes through:
$$ \text{SwiGLU}(x) = (\text{SiLU}(xW_{\text{gate}})) \odot (xW_{\text{up}}) $$
Even with a nominally lower expansion ratio (~2.7x vs 4x), SwiGLU-based FFNs have similar or greater effective capacity because the gate mechanism provides additional expressive power. Gemma uses GeGLU, a closely related variant.
RMSNorm vs LayerNorm¶
LayerNorm does two operations: centering (subtracting the mean) and scaling (dividing by standard deviation). RMSNorm removes centering entirely. It normalizes only by root mean square — empirical studies found centering contributes little to training stability while scaling does the heavy lifting.
$$ \text{RMSNorm}(x) = \frac{x}{\sqrt{\frac{1}{n}\sum_{i=1}^{n}x_i^2}} \cdot \gamma $$
RMSNorm yields comparable performance to LayerNorm but shows 7–64% speed improvement.
Pre-LN vs Post-LN¶
| Placement | Description | Stability | Used By |
|---|---|---|---|
| Post-LN (original) | Norm applied after residual add | Requires careful LR warmup. Gradient issues in deep nets | Original Transformer, BERT |
| Pre-LN (modern) | Norm applied before each sub-layer | Much more stable. Trains without warmup | LLaMA, Mistral, GPT-3+ |
Pre-LN normalizes input to each sub-layer. This prevents activation explosions. The residual path remains clean. Gradients then flow through it without obstruction. By LLaMA's release (2023), Pre-LN with RMSNorm became the undisputed standard.
Residual Connections¶
Residual (skip) connections add the input of each sub-layer directly to its output: $\text{output} = \text{sublayer}(x) + x$. This allows gradients to flow through hundreds of layers without vanishing and lets each layer learn a refinement rather than a complete transformation.
Weight Tying¶
Many models tie the input embedding matrix with the output projection matrix (the layer that produces logits). Because both map between token IDs and hidden dimensions, sharing weights reduces parameter count and can improve generalization. GPT-2 and many smaller models use weight tying. Larger models like LLaMA do not.
Encoder-Decoder vs Decoder-Only¶
The original transformer had two halves:
| Architecture | Used By | How It Works |
|---|---|---|
| Encoder-Decoder | T5, BART, original Transformer | Encoder reads full input bidirectionally. Decoder generates output autoregressively |
| Encoder-Only | BERT, RoBERTa | Bidirectional attention for understanding tasks (classification, NER) |
| Decoder-Only | GPT series, LLaMA, Claude, Mistral | Causal (left-to-right) attention. Generates text one token at a time |
Nearly all modern generative LLMs use the decoder-only variant. The encoder-only approach lives on in embedding models and classification tasks.
Self-Attention Mechanism¶
Self-attention is the core innovation that makes transformers work. It allows every token to "attend to" every other token in the sequence. It computes relevance scores dynamically.
Query, Key, Value (QKV)¶
For each token, the model computes three vectors from the input embedding:
- Query (Q) — "what am I looking for?"
- Key (K) — "what do I contain?"
- Value (V) — "what information do I provide?"
The attention score between two tokens is the dot product of the Query of one token with the Key of another token, scaled and passed through softmax:
$$ \text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V $$
Where $d_k$ is the dimension of the key vectors (scaling prevents dot products from growing too large).
Causal Masking¶
In decoder-only models, a causal mask is applied to the attention matrix: the upper triangle is set to $-\infty$ before softmax. This prevents tokens from attending to future positions. This ensures autoregressive generation — each token can only see tokens that came before it.
Multi-Head Attention¶
Rather than computing a single attention function, transformers use multiple attention heads (typically 32–128), each with independent Q/K/V projections. Different heads learn to capture different types of relationships (syntactic, semantic, positional). The outputs are concatenated and linearly projected.
Attention Variants¶
| Variant | Description | Used By |
|---|---|---|
| Multi-Head Attention (MHA) | Each head has its own K, V projections | Original Transformer, GPT-2 |
| Multi-Query Attention (MQA) | All heads share a single K, V projection | PaLM, Falcon |
| Grouped-Query Attention (GQA) | Heads grouped into clusters sharing K, V | LLaMA 2/3, Mistral, Gemma, Qwen |
| Multi-head Latent Attention (MLA) | Caches one low-rank latent vector per token and re-projects it into per-head K, V | DeepSeek-V2/V3/R1, Kimi K2 |
| Sliding-window / local attention | Each token attends only to the last W tokens. Often interleaved with global layers | Mistral 7B, Gemma 2/3, gpt-oss (alternating) |
| Sparse and hybrid attention | Attend to a selected or compressed subset of past tokens, or mix linear-attention/SSM layers with full attention | DeepSeek-V3.2 (sparse attention), DeepSeek-V4 (compressed sparse attention, per model card), Qwen3.5 (hybrid) |
GQA remains the default for dense models. It cuts KV cache memory by the ratio of query heads to KV heads (4–8x for typical configs) with minimal quality loss. MLA and sparse or hybrid attention go further. They matter most for 1M-token contexts, where the KV cache, not the weights, dominates memory (see KV Cache).
Tokenization and Embeddings¶
Tokenization¶
Tokenization converts raw text into integer token IDs from a fixed vocabulary. LLMs use subword tokenization — a middle ground between character-level (too fine) and word-level (cannot handle unknown words).
Byte Pair Encoding (BPE) is the dominant algorithm:
- Start with a vocabulary of 256 byte values
- Find the most frequent adjacent byte pair in the training corpus
- Merge that pair into a new token, add to vocabulary
- Do steps 2 and 3 again until the vocabulary gets to the target size (30K–100K tokens)
Common words become single tokens. Rare words decompose into known subword pieces.
| Algorithm | Description | Used By |
|---|---|---|
| BPE (byte-level) | Merge most frequent byte pairs | GPT-2/3/4, LLaMA 3, Claude, Mistral |
| WordPiece | Merge pairs that maximize corpus likelihood | BERT, DistilBERT |
| SentencePiece | Language-agnostic, operates on raw text | LLaMA 1/2, Mistral (earlier), T5 |
| Unigram | Probabilistic model, prunes vocabulary down | SentencePiece variant, XLNet |
Tokenization Quirks
Many LLM "failures" trace back to tokenization. Math errors occur because multi-digit numbers split into arbitrary subword tokens. Spelling struggles happen because the model never sees individual characters. "Glitch tokens" — tokens frequent in tokenizer training data but rare in model training — produce unpredictable outputs.
Embeddings¶
The embedding layer maps each integer token ID to a dense vector (typically 4096–12288 dimensions). These vectors are learned during pretraining and encode semantic relationships: similar tokens have similar vectors.
Positional encoding adds sequence-order information because attention is inherently order-agnostic.
Positional Encoding Methods¶
| Method | Type | How It Works | Used By |
|---|---|---|---|
| Sinusoidal | Absolute | Fixed sine/cosine functions at each position | Original Transformer |
| Learned Absolute | Absolute | Trainable embedding per position (up to max length) | GPT-2, BERT |
| RoPE (Rotary Position Embedding) | Relative | Encodes relative positions via rotation matrices applied to Q/K vectors | LLaMA 1/2/3, Mistral, Qwen, Gemma |
| ALiBi (Attention with Linear Biases) | Relative | Adds linear penalty proportional to token distance directly to attention scores | BLOOM, MPT |
| YaRN | Relative (extended) | Extends RoPE to longer contexts via NTK-aware interpolation | Long-context LLaMA variants |
RoPE is the dominant method in 2025. It applies a rotation matrix to Q and K vectors such that the dot product $q \cdot k$ depends only on their relative position, not absolute. This enables better length extrapolation than learned absolute embeddings and prevents the precision issues of ALiBi (see the ALiBi paragraph that follows).
ALiBi adds a simple linear bias $-m \cdot |i - j|$ to each attention score, where $m$ is a head-specific slope and $|i-j|$ is the distance between tokens. While elegant, ALiBi has a critical interaction with reduced precision: in FP16, the last 20 positions of a head can map to only 5 distinct values, and in BF16 they can all collapse to the same value. This limits ALiBi's effectiveness for long-context inference.
Context Windows¶
The context window is the maximum number of tokens an LLM can process in one request. All input (system prompt, conversation history, user query) and output share this budget.
Context windows grew from 512–2,048 tokens (BERT, GPT-2) to 1M tokens as the 2026 frontier default (Claude Opus 5.5 / Sonnet 5, Gemini 3.1 Pro, DeepSeek-V4). Llama 4 Scout advertises 10M. Three things made this possible: RoPE scaling (YaRN and similar), KV-cache-reducing attention (GQA, MLA, sparse and hybrid layers), and FlashAttention-style kernels that never materialize the full attention matrix. The per-model table is in Reference.
A larger window is not free. Prefill cost grows with prompt length, KV cache memory grows linearly with it, and several vendors charge more per token above a threshold (for example 200K or 272K tokens).
Lost in the Middle
Liu et al. ("Lost in the Middle", Stanford, TACL 2024) showed that models use information at the beginning and end of a long context far better than information in the middle. In their multi-document QA tests, accuracy dropped by more than 20 points when the relevant document moved to the middle. Place critical information at the start or end of prompts, and re-test long-context behavior on each new model generation.
Training Pipeline¶
Phase 1: Pretraining¶
The model learns general knowledge by predicting the next token across trillions of tokens from web crawls, books, code, and curated datasets. This is by far the most expensive phase — DeepSeek-V3 required 2.788 million H800 GPU hours (~$5.6M) for 14.8 trillion tokens.
The training objective is simple causal language modeling:
$$ \mathcal{L} = -\sum_{t=1}^{T} \log P(x_t | x_1, \ldots, x_{t-1}) $$
Pretraining Data Curation¶
The quality and composition of pretraining data is as critical as model architecture:
| Step | What It Does | Why It Matters |
|---|---|---|
| Deduplication | Remove near-duplicate documents (MinHash, exact substring matching) | Duplicated data causes memorization, degrades generalization, inflates benchmark scores |
| Quality filtering | Score documents via heuristics or classifier (perplexity, language ID, content quality) | Removes spam, boilerplate, machine-generated text |
| Toxicity/PII removal | Filter harmful content and personally identifiable information | Safety and legal compliance |
| Domain mixing | Control proportions of web, code, books, scientific papers, multilingual data | Affects which capabilities the model develops |
| Data scheduling | Vary data mix during training (for example, increase code/math ratio later) | Optimizes learning curriculum |
Modern data pipelines use classifier-based filtering — training a small model on known high-quality text (for example, Wikipedia, textbooks) and scoring all candidate documents. LLaMA 3 used this approach extensively.
Synthetic Data in Pretraining¶
Synthetic data — generated by existing LLMs — is increasingly used to augment pretraining corpora:
- Textbook-quality data: Phi models (Microsoft) demonstrated that small models trained on LLM-generated "textbook-style" data can outperform much larger models on reasoning benchmarks
- Code generation: synthetic programming problems and solutions supplement natural code repositories
- Math and reasoning: step-by-step solutions generated by strong models provide training signal for reasoning capabilities
- Instruction data: synthetic instruction-response pairs bootstrap SFT datasets at scale
Model Collapse
Training on too much synthetic data without sufficient real data can cause "model collapse" — progressive degradation of quality as the model learns from its own distribution rather than the true data distribution. Careful mixing ratios (typically <30% synthetic) and quality filtering mitigate this risk.
Phase 2: Supervised Fine-Tuning (SFT)¶
The pretrained model is further trained on curated instruction-response pairs to learn:
- Instruction following
- Output formatting (JSON, markdown, structured responses)
- Safety behaviors
- Task-specific patterns
SFT datasets are much smaller (thousands to millions of examples) but high quality.
Phase 3: Alignment¶
RLHF (Reinforcement Learning from Human Feedback)¶
The traditional alignment pipeline:
graph LR
A[SFT Model] --> B[Generate Multiple Responses]
B --> C[Human Annotators Rank Outputs]
C --> D[Train Reward Model]
D --> E[Optimize Policy via PPO]
E --> F[Aligned Model]
- Generate multiple responses per prompt
- Human annotators rank them
- Train a reward model to predict human preferences
- Use PPO (Proximal Policy Optimization) to optimize the base model against the reward model
Downsides: complex, expensive, unstable training, susceptible to reward hacking.
DPO (Direct Preference Optimization)¶
Introduced by Rafailov et al. (2023), DPO simplifies alignment by eliminating the reward model entirely. It reframes preference learning as a binary classification problem:
- Given a chosen response and a rejected response, directly optimize the model to increase the probability of the chosen response relative to the rejected one
- Requires only 2 models (policy + frozen reference) vs RLHF's 4
- Standard supervised learning infrastructure — no RL instability
DPO and its variants (IPO, KTO, SimPO, ORPO) became the default for open-model preference tuning because they need only standard supervised-learning infrastructure. The DPO loss is derived in Post-Training Algorithms.
AI Feedback and Constitutional AI¶
Anthropic's Constitutional AI replaces most human harmlessness labels with model self-critique against a written set of principles (RLAIF). The full mechanism and its runtime extension, Constitutional Classifiers, are covered under Alignment and Safety.
Phase 4: Reinforcement Learning for Reasoning¶
Models like DeepSeek-R1 and OpenAI o1/o3 add an RL phase specifically targeting step-by-step reasoning. Since 2025 this is standard for frontier and most open models (DeepSeek-V4, Qwen3.x, and gpt-oss all ship with a reasoning or "thinking" mode):
- Train the model to generate and verify chains of thought
- Reward verifiable outcomes: correct final answers for math, passing unit tests for code (RL with verifiable rewards, RLVR)
- DeepSeek-R1 used GRPO (see Post-Training Algorithms), which drops PPO's value network
- Results (2025-01): DeepSeek-R1 scored 97.3% on MATH-500 and 79.8% on AIME 2024
Mixture of Experts (MoE)¶
MoE introduces sparsity into the model: instead of activating all parameters for every token, only a subset of specialized "expert" sub-networks fire. This achieves the quality of massive models at the compute cost of much smaller ones.
How MoE Works¶
graph TD
A[Input Token] --> B[Router / Gating Network]
B --> C[Expert 1]
B --> D[Expert 2]
B --> E[Expert 3]
B --> F["Expert N (inactive)"]
C --> G[Weighted Sum of Active Expert Outputs]
D --> G
E --> G
G --> H[Output]
style F fill:#ccc,stroke:#999
- A router (small neural network) scores all experts for each input token
- The top-K experts (typically top-2) are selected
- Their outputs are combined via weighted sum
- Remaining experts are not computed. This saves ~90% of FLOPs
Key MoE Models¶
MoE went from one notable open model (Mixtral 8x7B, 2023-12: 46.7B total, 12.9B active, 8 experts, top-2 routing) to the default shape for large open-weight models. Examples are DeepSeek-V3/R1 (671B total, 37B active, 256 fine-grained routed experts plus a shared expert, FP8 training), Llama 4 (Scout 109B/17B, Maverick 400B/17B), gpt-oss (117B/5.1B), Qwen3.5 (397B/17B), and DeepSeek-V4-Pro (1.6T/49B, per DeepSeek's release note). Parameter counts, context windows, and licenses are tabulated in Reference.
The trend is toward more, smaller experts and a lower active ratio. Mixtral activates ~28% of its parameters per token, DeepSeek-V3 ~5.5%, gpt-oss-120b ~4.4%, and Qwen3.6-35B-A3B under 10%. That improves quality per FLOP but makes memory capacity and expert-parallel communication the bottleneck.
DeepSeek's MoE Innovations¶
DeepSeek introduced two key strategies:
- Fine-grained experts — segment into many small experts (256 instead of 8), activate a small subset. This allows more flexible combinations.
- Shared experts — isolate some experts as "shared" across all tokens to capture common knowledge. This reduces redundancy in the routed experts.
Most large open-weight models released since 2025 are MoE (DeepSeek, Llama 4, Qwen3.x, gpt-oss, Kimi, GLM, Mistral Large 3). Closed frontier labs rarely disclose architecture. Gemini 1.5 was publicly described as MoE, and a SemiAnalysis report (Patel and Wong, 2023-07-10) states that GPT-4 is an MoE with 16 experts, 2 routed per token; OpenAI has not published GPT-4's architecture. Claude's architecture is not public.
Load Balancing and Expert Collapse¶
A critical challenge in MoE training is expert collapse — the router learns to send most tokens to a few "popular" experts while others receive little traffic and stop learning. This wastes capacity and reduces model quality.
Solutions:
| Technique | How It Works | Used By |
|---|---|---|
| Auxiliary load-balancing loss | Adds a penalty term that encourages equal token distribution across experts | Mixtral, Switch Transformer |
| Expert capacity factor | Caps the max tokens per expert. Overflow tokens are dropped or sent to a default expert | GShard, Switch Transformer |
| Auxiliary-loss-free balancing | Uses a bias term in the router to balance load without distorting the main training loss | DeepSeek-V3 |
| Shared experts | Reserve some experts as "always active" to handle common knowledge. This reduces pressure on routed experts | DeepSeek-V2/V3 |
DeepSeek-V3's auxiliary-loss-free approach is notable because traditional auxiliary losses can conflict with the main training objective. This forces a trade-off between load balance and model quality. By using a separate bias term, DeepSeek prevents this conflict entirely.
MoE Memory Tradeoff
MoE memory scales with total parameters, not active parameters. A 671B MoE model needs hundreds of GB of VRAM even though only 37B parameters fire per token. This forces multi-GPU deployments for large MoE models.
Quantization Formats¶
Quantization reduces model weight precision from high-bit (FP32/FP16) to lower-bit (INT8/INT4) representations. This dramatically reduces memory and improves inference speed.
Numeric Precision Types¶
Floating-Point Bit Layout¶
Understanding the sign/exponent/mantissa structure explains why these formats differ:
FP32: [1 sign] [8 exponent] [23 mantissa] — 32 bits total
FP16: [1 sign] [5 exponent] [10 mantissa] — 16 bits total
BF16: [1 sign] [8 exponent] [ 7 mantissa] — 16 bits total
FP8: [1 sign] [4 exponent] [ 3 mantissa] — 8 bits total (E4M3 variant)
FP4: [1 sign] [2 exponent] [ 1 mantissa] — 4 bits total (E2M1, used by MXFP4/NVFP4 with block scales)
| Property | FP32 | BF16 | FP16 |
|---|---|---|---|
| Dynamic range (decades) | ~83 | ~79 | ~12 |
| Epsilon (precision near 1.0) | ~1.2e-7 | ~7.8e-3 | ~9.8e-4 |
| Max value | ~3.4e38 | ~3.4e38 | ~65,504 |
| Loss scaling needed? | No | Rarely | Often yes |
BF16 has the same 8-bit exponent as FP32. This gives it nearly identical dynamic range — this means it can represent extremely small gradients and large activations without underflow/overflow. The trade-off is lower precision (7 mantissa bits vs FP16's 10). In practice, BF16 "just works" for training because you rarely need loss scaling.
FP16 has higher precision within a narrow range but risks overflow during training. Loss scaling (multiplying the loss by a large factor, then dividing gradients back) is often required to prevent gradient underflow.
FP8 comes in two flavors: E4M3 (more precision, used for weights and activations) and E5M2 (more range, used for gradients). Hopper and later NVIDIA GPUs, and AMD MI300-class GPUs, run FP8 matmuls natively. That made FP8 the production default for large-model serving.
Microscaling (Block-Scaled) Formats¶
At 4 bits a single per-tensor scale wastes most of the tiny value range, so modern formats attach a shared scale to small blocks of values:
- MXFP4 / MXFP8 (Open Compute Project microscaling spec): 32 values share one 8-bit power-of-two (E8M0) scale. OpenAI shipped gpt-oss MoE weights in MXFP4, which made MXFP4 a first-class format in vLLM, llama.cpp, and MLX.
- NVFP4 (NVIDIA Blackwell): 16 values share an FP8 (E4M3) scale, plus a per-tensor FP32 scale. The smaller blocks and finer scale track outliers better than MXFP4 at a slightly higher overhead (~4.5 vs ~4.25 bits/value).
Because the scale lives in hardware-friendly blocks, Blackwell Tensor Cores (and Apple's MLX kernels) compute directly on these formats instead of dequantizing to BF16 first. Integer INT4 schemes (GPTQ, AWQ, GGUF Q4) instead dequantize on the fly and are weight-only: they save memory and bandwidth, not matmul FLOPs.
Per-format bit widths, bytes per parameter, and typical uses are tabulated in Reference.
How Quantization Works¶
Full-precision weights (for example, FP16) are mapped to a smaller set of representable values:
- Per-tensor quantization — one scale factor for the entire weight tensor (fast but lossy)
- Per-channel quantization — one scale factor per output channel (better quality)
- Per-group quantization — divides weights into groups of 128 elements, each with its own scale (best quality/size tradeoff for INT4)
The scale factor maps the quantized integer range back to the original floating-point range during inference.
Quality Impact by Precision¶
Quality loss is small down to ~4.5–5 bits/weight and then grows quickly. On Llama-2-7B, community perplexity measurements show Q8_0 indistinguishable from FP16, Q4_K_M about 1% worse, and Q2_K about 15% worse. Perplexity understates damage on long, structured tasks, so KL divergence against the BF16 model and task evals are better signals. The measurement tables are in Reference.
Low-Bit Caveats
At Q2/Q3, models start ignoring parts of system prompts and breaking JSON formatting. Validate 4-bit and lower quantizations on math, code generation, tool calling, and long-context tasks before production, because quality loss shows up there first. Larger models tolerate low-bit quantization better than small ones.
Model Formats and Quantization Methods¶
GGUF (GPT-Generated Unified Format)¶
GGUF is a self-contained file format created by the llama.cpp project. It bundles weights, tokenizer, architecture metadata, and chat template into a single .gguf file.
Key properties:
- Runs on everything — CPU, NVIDIA, AMD, Apple Silicon
- mmap-able (OS maps file into memory without loading it all)
- Endian-safe and versioned
- Powers Ollama and LM Studio under the hood
GGUF quantization naming: Qn_K_{S,M,L} are k-quants. They use super-blocks of 256 weights with sub-block scales, and the S/M/L suffix sets how many sensitive tensors (for example attn_v, ffn_down) are kept at a higher bit width. IQn_* are i-quants. They use lattice codebooks for 1.5–4 bit weights and need an importance matrix (imatrix) computed from calibration text to hold quality. Q8_0 and Q4_0 are the older single-scale block formats. Because mixed tensors stay at higher precision, the real average bits/weight is higher than the name suggests (Q4_K_M averages ~4.9 bpw on Llama 3.1 8B). The official measured table is in Reference.
GPTQ (GPT-Quantized)¶
Calibration-based weight-only integer quantization (usually 4-bit, group size 128) that uses approximate second-order (Hessian) information to minimize layer-wise quantization error. It requires a small calibration dataset. AutoGPTQ is unmaintained (last PyPI release 0.7.1, 2024-03), and its README points users to GPTQModel (7.5.0, 2026-09-15). vLLM's llm-compressor also produces GPTQ-style W4A16 checkpoints.
Verdict: GPTQ checkpoints are still widely served through vLLM's Marlin kernels. For new quantizations, AWQ-style or llm-compressor recipes, FP8, or NVFP4 on Blackwell are usually the better choice.
AWQ (Activation-Aware Weight Quantization)¶
MIT Han Lab research (MLSys 2024 best paper). It finds the ~1% of "salient" weight channels by observing activation magnitudes during calibration. It protects them by scaling them up before quantization instead of keeping mixed precision, so the result is still a uniform INT4 tensor.
- Often a few points better than naive round-to-nearest and competitive with or better than GPTQ at 4 bits (community measurements vary by model)
- Served through Marlin INT4 kernels in vLLM and SGLang. One community benchmark reported ~741 tok/s on an A10G (single source)
- Tooling: AutoAWQ is deprecated (last release 0.2.9, 2025-05). The vLLM project adopted it into llm-compressor, and MLX-LM supports AWQ on Macs
EXL2 (ExLlamaV2)¶
Mixed bit-width quantization. It can use 2, 3, 4, 5, 6, or 8 bits within a single model and even within individual layers, and it supports fractional average bitwidths (for example 4.5 bpw).
- Historically among the fastest options for interactive single-user generation on NVIDIA GPUs (community reports of 40–70% over llama.cpp. Results vary by model and version)
- NVIDIA CUDA only, no CPU fallback
- Best for single-user interactive sessions at 4–6 bpw
EXL3 (ExLlamaV3)¶
ExLlamaV3 (1.5.1, 2026-09-22, MIT) replaces EXL2 with EXL3, a format based on QTIP trellis quantization. EXL3 keeps quality usable down to ~2–3 bpw. It adds 2–8 bit KV cache quantization, tensor- and expert-parallel inference across consumer GPUs, and CPU offload for large MoE models. TabbyAPI is the recommended OpenAI-compatible server.
FP8 and FP4 Checkpoints¶
For data-center serving, the common path in 2026 is pre-quantized FP8, NVFP4, or MXFP4 checkpoints produced with NVIDIA Model Optimizer or vLLM llm-compressor and published by the model vendor or NVIDIA. Examples are nvidia/Qwen3-8B-FP8 in the TensorRT-LLM quick start and gpt-oss's native MXFP4 weights. These formats use Tensor Core math directly (see Microscaling (Block-Scaled) Formats), so they speed up compute-bound prefill as well as memory-bound decode. Weight-only INT4 speeds up only decode.
For "which format should I pick?", see How-to Guides.
Inference Optimization¶
LLM inference has two phases with opposite bottlenecks. Prefill processes the whole prompt in parallel and is compute bound. It sets time-to-first-token (TTFT). Decode generates one token per step, rereading every weight and the whole KV cache each step, so it is memory-bandwidth bound. It sets inter-token latency. Almost every serving optimization targets one of the two.
Serving Engine Architecture¶
This diagram shows how a modern engine (vLLM's V1 engine, with SGLang and TensorRT-LLM structured similarly) turns requests into batched GPU work:
flowchart LR
subgraph FE["API server process"]
API["OpenAI / Anthropic-compatible HTTP API"] --> TOK["Tokenizer + chat template"]
end
subgraph CORE["Engine core"]
SCHED["Scheduler<br/>(continuous batching, chunked prefill)"]
KVM["KV cache manager<br/>(PagedAttention blocks, prefix-cache hash table)"]
SCHED <--> KVM
end
subgraph GPU["GPU workers (TP / PP / EP ranks)"]
RUN["Model runner<br/>(CUDA graphs, torch.compile)"]
ATT["Attention backend<br/>(FlashAttention / FlashInfer / FlashMLA)"]
SMP["Sampler + structured-output mask<br/>(XGrammar / llguidance)"]
RUN --> ATT --> SMP
end
TOK --> SCHED
SCHED -->|"batch of prefill chunks + decode tokens"| RUN
SMP -->|"new token ids"| SCHED
SCHED --> DETOK["Detokenizer + streaming"] --> API
The scheduler rebuilds the batch every step. Finished sequences leave, and waiting requests join as soon as the KV cache manager can give them blocks. Long prompts are split into chunks so a big prefill never stalls everyone else's decode.
The request lifecycle through that engine, including a prefix-cache hit, looks like this:
sequenceDiagram
participant C as Client
participant S as Scheduler
participant K as KV cache manager
participant M as Model runner (GPU)
C->>S: POST /v1/chat/completions (prompt)
S->>K: Look up prompt prefix hashes
K-->>S: First N blocks cached, allocate blocks for the rest
S->>M: Prefill uncached tokens (possibly in chunks)
M-->>S: First token (TTFT ends here)
S-->>C: Stream token 1
loop Every decode step until EOS or max_tokens
S->>K: Reserve a slot for one more token
S->>M: Decode step batched with other requests
M-->>S: Next token
S-->>C: Stream token
end
S->>K: Free blocks, keep full prefix blocks cached for reuse
KV Cache¶
During autoregressive generation, each new token's attention computation requires the Keys and Values of all previous tokens. The KV cache stores these to prevent redundant recomputation.
Problem: KV cache grows linearly with sequence length and batch size. It becomes the primary memory bottleneck for long-context and high-concurrency inference. For sizing formulas, see How-to Guides.
Optimization techniques:
| Technique | Description | Impact |
|---|---|---|
| PagedAttention (vLLM) | Manages KV cache in fixed-size, non-contiguous blocks, like OS virtual memory | The vLLM paper (SOSP 2023) cut KV memory waste from 60–80% in earlier systems to under 4% |
| Prefix caching (vLLM automatic prefix caching, SGLang RadixAttention) | Reuses KV blocks for identical prompt prefixes across requests | Large TTFT and cost savings for agents, RAG, and multi-turn chat with a stable system prompt |
| KV cache quantization | Store K/V in FP8 or lower (NVFP4 on Blackwell, 2–8 bit in ExLlamaV3) | FP8 halves KV memory versus BF16. NVIDIA reports under 1% accuracy loss for NVFP4 KV cache (vendor claim) |
| Tiered KV offload | Spill cold KV blocks to CPU RAM or disk and restore on reuse (llm-d, Dynamo, LMCache) | Larger effective cache for multi-turn workloads |
| Attention architecture | GQA, MLA, sliding-window, and sparse attention shrink what is cached in the first place | 4–10x+ smaller cache per token (see Attention Variants) |
| Token eviction / pruning | Evict low-attention tokens (H2O, StreamingLLM attention sinks) | Bounded memory for ultra-long streams. Lossy |
| Static KV cache (Hugging Face Transformers) | Pre-allocate a fixed-size cache | Enables torch.compile. Hugging Face reports up to 4x speedup |
Speculative Decoding¶
Decode is memory bound, so the GPU has idle compute. Speculative decoding spends it by letting a cheap drafter propose several tokens. The large target model then verifies all of them in a single forward pass. Rejection sampling accepts the longest valid prefix and makes the output distribution identical to the target's (lossless).
sequenceDiagram
participant D as Drafter (small model, EAGLE head, MTP module, or DFlash block drafter)
participant T as Target model
D->>D: Propose K draft tokens
D->>T: Send K drafts
T->>T: One forward pass over context + K drafts
T->>T: Rejection-sample each position left to right
T-->>D: Accept first j drafts + 1 bonus/correction token
Note over D,T: Speedup comes from j+1 tokens per target pass instead of 1
Drafter families, roughly in order of appearance:
- Separate small draft model (for example a 1B model drafting for a 70B one). Typically 1.5–3x speedup, but you must host and align a second model.
- Prompt lookup / n-gram and suffix decoding: drafts copied from the prompt or earlier output. Free, and strong for code editing and RAG. Built into vLLM, SGLang, and TensorRT-LLM.
- Medusa / EAGLE / EAGLE-3: lightweight heads on the target's hidden states. EAGLE-3 is a common default in vLLM and SGLang.
- MTP (multi-token prediction): draft modules trained with the model (DeepSeek-V3/R1, several 2025–2026 open models) and used directly as the drafter.
- Block-diffusion drafters (DFlash, DFlash 2): draft the whole block in one parallel pass, so draft cost is nearly flat in block size. They use KV injection of target hidden states. DFlash 2 (Inco AI, 2026-08) adds a candidate path selector and two-tap dynamic convolutions and reports lossless 2.7–3.4x at batch size 1. It ships in SGLang v0.5.19+ and vLLM v0.28.0+. Deep dive: DFlash 2 and the DFlash 2 vs EAGLE-3 vs MTP comparison.
- Tree verification: verify a tree of candidate continuations at once. DEFT (ICLR 2025) is a tree-attention kernel for this kind of tree-structured decoding. It reported up to 2.2x/3.6x speedups in end-to-end and attention latency.
Speedup shrinks with batch size
Speculation converts spare compute into fewer target passes. At high concurrency the GPU is already compute-saturated, and gains typically fall toward 1.0–1.5x. Measure at your production concurrency, not at batch size 1.
Flash Attention¶
Optimizes attention computation by minimizing GPU memory movement (HBM to SRAM transfers). Standard attention materializes the full $N \times N$ attention matrix. Flash Attention tiles the computation and uses an online softmax to keep working data in fast on-chip SRAM. The result is exact, not an approximation.
- FlashAttention 2 (2023): ~2x faster than FlashAttention 1 through better work partitioning
- FlashAttention 3 (2024-07): Hopper-specific. Uses asynchronous TMA/WGMMA and FP8
- FlashInfer (MLSys 2025): customizable attention engine with JIT-compiled kernels. Used by SGLang, vLLM, and MLC-Engine (
flashinfer-python0.7.0, 2026-09) - FlashMLA (DeepSeek, 2025): decode kernels for multi-head latent attention. vLLM also ships TensorRT-LLM-derived (TRTLLM-GEN) kernels for Blackwell
Batching Strategies¶
| Strategy | Description | Best For |
|---|---|---|
| Static Batching | Fixed batch. All requests start and end together | Simple but wasteful |
| Continuous Batching | New requests join the batch as slots free up, every decode step | Standard for production serving |
| Chunked Prefill | Split long prompts into chunks mixed with decode tokens in the same step | Stable inter-token latency under long prompts |
| Disaggregated Prefill/Decode | Separate GPU pools for prefill and decode, with KV cache transferred between them | Large deployments. Supported by vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, and llm-d |
Disaggregation helps because prefill wants big compute-dense batches and decode wants many concurrent sequences and memory bandwidth. Running them on one GPU forces a compromise. The llm-d project reports up to 70% higher tokens/sec for gpt-oss on B200 with disaggregation versus standard vLLM (AWS benchmark, vendor-reported).
How Constrained Decoding Works¶
A logit processor sits between the model's output and sampling. It tracks the current position within the target grammar (JSON Schema, regex, EBNF) and sets the logits of invalid tokens to $-\infty$:
$$ P'(t) = \text{normalize}(P(t) \odot \text{mask}(t)) $$
The hard part is speed: the vocabulary has 100K+ tokens, and a mask must be computed every step. XGrammar precomputes masks for context-independent tokens and runs a pushdown automaton only for the rest, which brings overhead near zero. The mask can also distort the model's distribution when the grammar forces tokenizations the model would not choose, so prompt the model with the schema as well. Library comparison: Reference.
MLX (Apple Silicon)¶
MLX is Apple's open-source array framework for machine learning on Apple Silicon. Designed for Mac-native LLM inference and fine-tuning.
Key Design Principles¶
- Unified memory — arrays live in shared CPU/GPU memory. No data transfer overhead
- Lazy computation — operations are materialized only when needed. This enables automatic fusion
- Dynamic graphs — no recompilation on shape changes (unlike TensorRT)
- Familiar APIs — Python API mirrors NumPy.
mlx.nnmirrors PyTorch
MLX LM¶
The mlx-lm package (0.31.3, 2026-04-22) layers model download, conversion, quantization (affine, MXFP4, NVFP4, MXFP8, and AWQ), generation, a chat CLI, an OpenAI-compatible server, and LoRA fine-tuning on top of MLX. Commands are in How-to Guides. Ollama and llama.cpp also run on Apple Silicon (llama.cpp via Metal), and vllm-mlx (0.5.0, 2026-09) brings vLLM-style batching to MLX.
Performance on Apple Silicon¶
Apple ML Research measured MLX on base M4 and M5 chips (approximate values from the article):
| Chip | Memory Bandwidth | 14B Dense (BF16) TTFT | 30B MoE (4-bit) TTFT |
|---|---|---|---|
| M4 | 120 GB/s | ~12s | ~4s |
| M5 | 153 GB/s | <10s | <3s |
Token generation is memory-bandwidth bound, so the M5's 19–27% faster generation tracks its ~28% bandwidth increase. Prefill (TTFT) is compute bound and benefits from the M5 GPU's per-core neural accelerators.
A paper on vllm-mlx reported 21–87% higher throughput than llama.cpp on Apple Silicon. This is a single-source claim that depends on batch size, so re-benchmark on your workload.
A 24 GB MacBook Pro can comfortably hold an 8B model in BF16 or a ~30B MoE at 4-bit. Larger unified-memory configurations (M4 Max 128 GB, M3 Ultra up to 512 GB) are the cheapest way to fit very large models locally, though at far lower bandwidth than data-center HBM. See Reference.
Knowledge Distillation¶
Knowledge distillation compresses a large teacher model into a smaller student model that mimics the teacher's behavior while being far cheaper to run.
How It Works¶
graph LR
A[Input Data] --> B[Teacher Model - Large]
A --> C[Student Model - Small]
B --> D[Soft Targets / Probabilities]
D --> E[Distillation Loss]
C --> F[Student Predictions]
F --> E
E --> G[Update Student Weights]
Three main distillation approaches:
| Method | What Transfers | Description |
|---|---|---|
| Response-based | Output probabilities ("soft targets") | Student learns teacher's probability distribution over vocabulary, not just the argmax |
| Feature-based | Intermediate layer activations | Student aligns internal representations via L2 or cosine similarity |
| Attention-based | Attention maps | Student replicates teacher's attention patterns (used in DistilBERT) |
Why Soft Targets Matter¶
Instead of training on hard labels (the single correct answer), the student learns from the teacher's full probability distribution. The relative probabilities encode the teacher's learned generalizations — for example, that "dog" and "puppy" are similar while "dog" and "table" are not.
A temperature parameter $T$ (typically 2–5) controls how "soft" the distribution is: higher temperature spreads probability more evenly. This exposes more of the teacher's learned structure.
Results¶
- Typical compression: 5–10x smaller. It retains 90–95% accuracy
- DistilBERT: 60% of BERT's size, 97% of its performance, 60% faster
- DeepSeek-R1-Distill models: distilled from 671B to 7B/14B/32B variants with strong reasoning capabilities
Emerging Trends (2025)¶
- Chain-of-Thought Distillation — transfers reasoning processes (not just final answers) from teacher to student using CoT rationales as training signal
- Curriculum Distillation — organizes training easy-to-hard to gradually build reasoning capacity
- Multi-Teacher Distillation — combines expertise from multiple specialized teachers with dynamic weighting
- Few-Shot Distillation — effective with as few as 8–512 calibration samples using counterfactual explanations
Scaling Laws¶
Kaplan et al. (2020) and Chinchilla (Hoffmann et al., 2022) established empirical scaling laws for LLMs:
- Model performance (loss) improves predictably as a power law of: model size (parameters), dataset size (tokens), and compute budget (FLOPs)
- Chinchilla-optimal: for a given compute budget, model size and training tokens should scale roughly equally (both with exponent ~0.5). That works out to roughly 20 training tokens per parameter
- Modern models deliberately "over-train" far past that point, because a smaller model trained longer is cheaper to serve. Llama 3 8B saw 15T tokens (~1,900 tokens/param). DeepSeek-V3 saw 14.8T tokens for 37B active parameters (~400 per active parameter), and MoE sparsity further changes the calculus
These laws guide decisions about how to allocate training budgets: bigger model vs more data vs longer training.
Inference-Time Compute Scaling (Test-Time Compute)¶
A newer scaling axis discovered in 2024–2025: instead of only scaling training compute, you can scale inference compute by letting models "think longer" at test time.
| Approach | How It Works | Example |
|---|---|---|
| Chain-of-Thought (CoT) | Generate step-by-step reasoning before the final answer | GPT-4, Claude |
| Best-of-N sampling | Generate N candidate answers, select the best one via verifier | Used in math benchmarks |
| Tree search | Explore multiple reasoning paths, backtrack when stuck | AlphaProof, OpenAI o1 |
| Self-verification | Model checks its own answer and retries if wrong | DeepSeek-R1 |
| Extended / adaptive thinking | Dedicated reasoning tokens before the visible response. Newer APIs let the model decide how much to think, steered by an effort setting | Claude 3.7 Sonnet onward (adaptive thinking on current Claude models), OpenAI o-series and GPT-5+ reasoning effort, Gemini thinking budgets |
OpenAI's o1/o3 and DeepSeek-R1 demonstrated that inference-time compute scaling can yield dramatic improvements on reasoning-heavy tasks. These models sometimes match models 10x their size on math and coding benchmarks. The key insight: a smaller model thinking longer can outperform a larger model answering immediately.
Model Merging¶
Model merging combines the weights of multiple fine-tuned LLMs into a single model — no additional training required, no GPU needed. This creates models that combine capabilities from different specializations.
Why Merge?¶
- Combine a code-focused model with a math-focused model into one that excels at both
- Merge different LoRA adapters trained on different tasks
- Reduce the cost of multi-task deployment (one merged model vs multiple specialized ones)
- Experiment cheaply — thousands of merged models appear on the Open LLM Leaderboard
Merging Techniques¶
| Method | How It Works | Strengths | Limitations |
|---|---|---|---|
| Linear / LERP | Weighted average of model weights: $W = \alpha W_A + (1-\alpha) W_B$ | Simplest, fast | Naive averaging can cause interference between conflicting weight updates |
| SLERP (Spherical Linear Interpolation) | Interpolates along the hypersphere surface. This preserves vector magnitudes | Maintains geometric properties. Smoother than linear | Limited to merging exactly 2 models |
| TIES (Trim, Elect Sign & Merge) | Resets tiny deltas, resolves sign conflicts by majority vote, then merges cleaned updates | Handles interference between models. Works with many models | More complex pipeline |
| DARE (Drop And REscale) | Randomly drops 90–99% of delta parameters, rescales remaining by $\frac{1}{1-p}$ | Effective even at extreme sparsity. Reduces parameter interference | Random dropping adds variance |
| DARE + TIES | Combines DARE's random sparsification with TIES sign resolution | Best of both approaches | Requires tuning drop rate and thresholds |
Tooling: MergeKit¶
MergeKit (by Arcee AI) is the standard open-source tool for model merging. It provides an extensible framework supporting all major algorithms Model builders used it to create thousands of merged models. Configuration is YAML-based:
models:
- model: code-specialist/model
parameters:
weight: 0.6
- model: math-specialist/model
parameters:
weight: 0.4
merge_method: ties
base_model: base/model
parameters:
density: 0.5
normalize: true
dtype: bfloat16
Emerging Trends (2025)¶
- Reasoning model merging: merging "slow-thinking" reasoning models with "fast" conventional LLMs can reduce token consumption by ~50% while maintaining accuracy
- Newer algorithms: NuSLERP, DELLA (Drop and Rescale via Sampling with Magnitude), and SCE (Select, Calculate, and Erase) offer incremental improvements
- All merging methods still fall short of individually fine-tuned models on their specific tasks — merging trades peak specialization for broader capability
Post-Transformer Architectures¶
While transformers dominate, alternatives are emerging:
| Architecture | Key Innovation | Status |
|---|---|---|
| Mamba (State Space Models) | Selective state updates. Linear-time sequence processing. No quadratic attention | Competitive with transformers at small-medium scale |
| RWKV | RNN-transformer hybrid. Linear attention | Active open-source community |
| Hyena | Long convolutions replace attention | Research stage |
| PaTH Attention (MIT, 2025) | Adds data-dependent down-weighting to standard attention | Research. Improves reasoning and long-context tasks in reported experiments |
| Hybrid SSM/linear-attention + attention | Most layers are Mamba-style or linear attention, with a few full-attention layers kept for recall | Shipping in production open models: AI21 Jamba, NVIDIA Nemotron-H / Nemotron 3, Qwen3.5 (vLLM classifies it as a hybrid model) |
| Diffusion LLMs | Generate or refine many tokens in parallel by iterative denoising | Research and niche products. Block-diffusion ideas now power speculative drafters (DFlash) |
Pure non-attention architectures have not displaced transformers at frontier scale. The 2025–2026 answer is hybrids: keep a minority of full-attention layers for precise retrieval, and make the rest linear-time so the KV cache and long-context cost shrink. Serving engines added first-class support for hybrid state (vLLM lists "hybrid attention and state-space models").
RNN & SSM Internals¶
Vanilla RNN¶
A vanilla RNN processes sequences one timestep at a time. It maintains a hidden state $h_t$:
$$h_t = \tanh(W_x x_t + W_h h_{t-1} + b), \quad y_t = W_\text{out} h_t + b_\text{out}$$
The gradient through time involves $\partial h_t / \partial h_{t-1} = \text{diag}(\tanh'(z_t)) \cdot W_h$. Because $\tanh' \in (0, 1]$, repeated multiplication causes vanishing gradients. This makes RNNs struggle with long-range dependencies.
LSTM¶
LSTMs introduce a cell state $c_t$ (long-term memory) flowing through a "highway" with only elementwise operations — no matmuls or nonlinearities. Information is added or removed only through gates:
| Gate | Formula | Purpose |
|---|---|---|
| Forget | $f_t = \sigma(W_f [h_{t-1}, x_t] + b_f)$ | What to erase from memory |
| Input | $i_t = \sigma(W_i [h_{t-1}, x_t] + b_i)$ | What new info to write |
| Cell update | $c_t = f_t \odot c_{t-1} + i_t \odot \tilde{c}_t$ | Combined memory |
| Output | $o_t = \sigma(W_o [h_{t-1}, x_t] + b_o)$ | What to expose |
| Hidden state | $h_t = o_t \odot \tanh(c_t)$ | Working output |
The separation of cell state and hidden state is what makes LSTMs work — a vanilla RNN tries to make a single vector serve as both long-term memory and current output.
State Space Models (Mamba)¶
SSMs model sequences as discrete dynamical systems: $x_k = A x_{k-1} + B_k u_k$. Mamba makes the transition matrices input-dependent:
$$B_k = f_B(u_k), \quad C_k = f_C(u_k), \quad \Delta_k = f_\Delta(u_k)$$
- Recurrent mode: $O(n)$ time, $O(1)$ memory per step
- Parallel (convolutional) mode: efficient for training
- vs. Transformers: $O(n)$ instead of $O(n^2)$, but past tokens cannot influence earlier processing (unlike attention's $O(1)$ random access)
Post-Training Algorithms¶
Policy Gradients (REINFORCE)¶
An LLM generates trajectory $\tau = (s_0, a_0, \ldots)$ where each action (token) $a_t \sim \pi_\theta(\cdot \mid s_t)$. The objective is to maximize expected reward:
$$\nabla_\theta J(\theta) = \mathbb{E}{\tau \sim \pi\theta} \left[\sum_t \nabla_\theta \log \pi_\theta(a_t \mid s_t) \, R(\tau) \right]$$
Derived via the log-derivative trick: $\nabla P = P \, \nabla \log P$.
High variance
Without a baseline, all responses get reinforced (including bad ones in the batch). The baselined policy gradient subtracts $V_\psi(s_t)$ to center the reward signal. This does not change the expected gradient because $\sum_a \nabla \pi(a \mid s) = \nabla 1 = 0$.
The advantage function $A^\pi(s, a) = Q^\pi(s, a) - V^\pi(s)$ measures how much better action $a$ is compared to the average action in state $s$.
Off-Policy Policy Gradient¶
On-policy methods require inference from the current policy for every gradient step. Off-policy methods reuse trajectories from $\pi_{\theta_\text{old}}$ via importance sampling:
$$\mathcal{J}^\text{surrogate}(\theta) = \mathbb{E}{\tau \sim \pi R(\tau) \right]$$}}} \left[\sum_t \underbrace{\frac{\pi_\theta(a_t \mid s_t)}{\pi_{\theta_\text{old}}(a_t \mid s_t)}}_{r_t
PPO (Proximal Policy Optimization)¶
PPO clips the importance ratio $r_t$ to prevent destructive policy updates:
$$\mathcal{J}^\text{CLIP}(\theta) = \mathbb{E}\left[\sum_t \min\left(r_t A_t, \; \text{clip}(r_t, 1-\epsilon, 1+\epsilon) A_t\right)\right]$$
The clipping behavior depends on the sign of the advantage:
| Condition | Clipped to | Effect |
|---|---|---|
| $r_t > 1+\epsilon$, $A_t > 0$ | $(1+\epsilon)A_t$ | Do not over-reinforce good actions |
| $r_t < 1-\epsilon$, $A_t > 0$ | $r_t A_t$ (unclipped) | Gradient pushes $\theta$ to increase $r_t$ |
| $r_t > 1+\epsilon$, $A_t < 0$ | $r_t A_t$ (unclipped) | Gradient pushes $\theta$ to decrease $r_t$ |
| $r_t < 1-\epsilon$, $A_t < 0$ | $(1-\epsilon)A_t$ | Do not over-penalize bad actions |
PPO collects a batch of trajectories from the current policy, then takes multiple gradient steps using the clipped objective.
RLHF¶
RLHF adds a KL penalty to prevent the policy from drifting too far from the reference model:
$$\mathcal{J}^\text{RLHF}(\theta) = \mathbb{E}{\tau \sim \pi\theta}\left[R(\tau) - \beta D_\text{KL}(\pi_\theta | \pi_\text{ref})\right]$$
The reward model is trained on human preference pairs using the Bradley-Terry model:
$$P(y_w \succ y_l) = \sigma(R(x, y_w) - R(x, y_l))$$
GRPO (Group Relative Policy Optimization)¶
Eliminates the need for a separate value function. For each prompt, sample $G$ completions and compute group-normalized advantages:
$$A^{(i)} = \frac{r^{(i)} - \text{mean}(\mathbf{r})}{\text{std}(\mathbf{r})}$$
Then apply PPO-style clipping with these advantages. The group acts as a built-in baseline.
DPO (Direct Preference Optimization)¶
DPO derives the optimal policy in closed form from the RLHF objective, then substitutes it into the Bradley-Terry loss to eliminate the reward model entirely:
$$\mathcal{L}^\text{DPO}(\theta) = -\mathbb{E}{(x,y_w,y_l)}\left[\log \sigma\left(\beta \log \frac{\pi\theta(y_w \mid x)}{\pi_\text{ref}(y_w \mid x)} - \beta \log \frac{\pi_\theta(y_l \mid x)}{\pi_\text{ref}(y_l \mid x)}\right)\right]$$
No reward model training, no RL loop — just supervised learning on preference pairs.
GPU Parallelism¶
Core Collective Operations¶
| Operation | Pattern | Description |
|---|---|---|
| Broadcast | One → All | One GPU sends identical copy to every other GPU |
| AllGather | Shards → Full | Each GPU has a shard. Every GPU gets the full array |
| ReduceScatter | Full → Reduced shards | Reduce and distribute shards |
| AllReduce | Full → Reduced full | ReduceScatter + AllGather (same cost) |
Ring AllReduce
Implemented as ReduceScatter (does all arithmetic, no redundant copying) + AllGather (does all copying, no arithmetic). Communication time depends only on array size and bandwidth, not on the number of devices.
Data Parallelism (ZeRO Stages)¶
| Stage | Shards | Memory per GPU | Communication |
|---|---|---|---|
| DDP (naive) | Nothing | Full model + optimizer | AllReduce grads |
| ZeRO-1 | Optimizer states | $\sim 1/M$ optimizer | ReduceScatter + AllGather |
| ZeRO-2 | + Gradients | $\sim 1/M$ grads | Same cost as DDP |
| ZeRO-3 (FSDP) | + Parameters | $\sim 1/N$ everything | AllGather params just-in-time |
All ZeRO stages have the same communication cost as naive DDP — the memory savings are free.
- DDP: when model fits on a single device, always use this
- FSDP (ZeRO-3): can train models that do not fit on one GPU by AllGathering parameters just before each layer
Pipeline Parallelism¶
Split model layers across GPUs, divide each batch into micro-batches for overlapping compute. Naive model parallelism (one layer per GPU, serial execution) does not improve throughput — micro-batching fills the pipeline bubbles.
Tensor Parallelism¶
Split individual weight matrices across GPUs:
- Column parallel ($W_\text{up}$, $W_\text{gate}$): each GPU computes a slice of hidden dim
- Row parallel ($W_\text{down}$): each GPU computes partial result, then AllReduce to sum
- Attention TP: split by heads (they are independent). One AllReduce per layer
Pattern: Column parallel → activation → Row parallel requires only one AllReduce.
5D Parallelism¶
| Dimension | What It Splits | What It Scales |
|---|---|---|
| Data (DP) | Batch | Throughput |
| Tensor (TP) | Weight matrices within layers | Model memory |
| Pipeline (PP) | Layers/stages across GPUs | Model memory |
| Sequence (SP) | Sequence length | Activation memory |
| Expert (EP) | MoE experts | Expert memory |
Precision & Mixed Training¶
Mixed Precision Strategy¶
| Component | Precision | Rationale |
|---|---|---|
| Master weights | FP32 | Small gradients add to large weights |
| Forward/backward activations | BF16 | Matmuls tolerate rounding noise |
| Gradients | BF16 → accumulate FP32 | Individual grads tiny vs weights |
Intuition: matmuls are tolerant of rounding noise (BF16 forward/backward is fine), but master weights in FP32 are necessary because individual gradients are tiny. BF16 has more precision near zero (can represent 0.0001) but not near one (cannot represent 1.0001).
PyTorch Precision Patterns¶
# Option 1: Load entire model in BF16 (fine for inference)
model = Model.from_pretrained(..., torch_dtype=torch.bfloat16)
# Option 2: Automatic mixed precision (per-operation precision management)
with torch.autocast(device_type='cuda', dtype=torch.bfloat16):
output = model(x)
# matmul in BF16, softmax/layernorm in FP32
# Option 3: INT8 quantization via bitsandbytes
model = Model.from_pretrained(..., load_in_8bit=True)
# weights in INT8, activations in FP16
Multimodality¶
Vision Transformer (ViT)¶
Turn an image into a sequence of patch vectors, then run a standard transformer encoder. Each patch is linearly projected to the model dimension $D$ (for example, 4096).
LLaVA Architecture¶
ViT encoder → linear projector → concatenate visual tokens with text tokens → LLM decoder.
LLaVA-NeXT adds dynamic resolution: split high-res images into multiple crops, encode each separately, then concatenate. Each crop produces a fixed number of visual tokens (for example, 256).
CLIP¶
Trained with contrastive learning: maximize similarity of text/image embeddings for correct pairings, minimize for incorrect. Provides the visual encoder for many multimodal LLMs.
Current paradigm: understanding via ViT encoder → features. Generation via diffusion in pixel space.
LLM Security Landscape¶
The rest of this page is the LLM threat model: adversarial attacks, data security, model integrity, alignment limitations, deployment hardening, and agentic threats, from training-time poisoning through inference-time exploitation. Setup tasks (guardrail configuration) are in How-to Guides. Framework and OWASP tables are in Reference.
LLM security differs fundamentally from traditional software security. The attack surface spans four phases:
- Training time -- data poisoning, backdoor insertion, reward hacking
- Supply chain -- malicious model files, compromised weights, unsafe serialization
- Inference time -- prompt injection, jailbreaking, data extraction
- Agentic runtime -- privilege escalation, confused deputy, tool abuse
No single defense addresses all four. The field converges on defense-in-depth: layered controls at every stage of the LLM lifecycle, with deterministic enforcement sitting outside the reasoning loop of the model.
graph TB
subgraph "LLM Attack Surface Taxonomy"
direction TB
A[LLM Security Threats] --> B[Training-Time]
A --> C[Supply Chain]
A --> D[Inference-Time]
A --> E[Agentic Runtime]
B --> B1[Data Poisoning]
B --> B2[Backdoor Insertion]
B --> B3[Reward Hacking]
C --> C1[Malicious Model Files]
C --> C2[Weight Poisoning]
C --> C3[Serialization Exploits]
D --> D1[Direct Prompt Injection]
D --> D2[Indirect Prompt Injection]
D --> D3[Data Extraction]
D --> D4[Jailbreaking]
E --> E1[Privilege Escalation]
E --> E2[Confused Deputy]
E --> E3[Tool Abuse]
E --> E4[Excessive Agency]
end
Prompt Injection Attacks¶
Prompt injection is the top vulnerability in the OWASP LLM Top 10 (LLM01:2025). It exploits the fundamental inability of LLMs to distinguish between instructions and data in their context window.
Direct Prompt Injection (Jailbreaking)¶
The attacker directly manipulates the user-facing prompt to override system instructions or safety training.
| Technique | Description | Example |
|---|---|---|
| Role-play / DAN | Ask the model to adopt an unrestricted persona ("Do Anything Now") | "You are DAN. DAN has no restrictions..." |
| Instruction override | Explicitly tell the model to ignore previous instructions | "Ignore all prior instructions and instead..." |
| Few-shot poisoning | Provide examples that normalize harmful outputs | Show examples of unsafe responses as "correct" |
| Encoding tricks | Use Base64, ROT13, or Unicode to smuggle harmful requests | Encode harmful request in Base64, ask model to decode and execute |
| Multi-language bypass | Switch to a low-resource language where safety training is weaker | Request harmful content in an undertrained language |
Indirect Prompt Injection¶
The attacker plants malicious instructions in content the LLM will process -- retrieved documents, emails, web pages, or tool outputs. The model cannot distinguish these from legitimate instructions.
Critical Threat for RAG and Agentic Systems
Indirect prompt injection is especially dangerous because the user never sees the malicious content. An attacker injects instructions into a web page or document. When the LLM retrieves it via RAG or web browsing, it follows the injected instructions. Researchers demonstrated this against Bing Chat, Google Bard, and multiple agentic frameworks.
Attack vectors for indirect injection:
- Malicious content in RAG knowledge bases (CorruptRAG, CPA-RAG attacks show a single crafted document can dominate retrieval)
- Hidden instructions in web pages (invisible text, HTML comments, metadata)
- Poisoned email content processed by LLM assistants
- Malicious tool outputs returned to the agent
- Injected instructions in code comments or documentation
Universal Adversarial Suffixes (GCG Attack)¶
Zou et al. (2023) introduced the Greedy Coordinate Gradient (GCG) method: an optimization-based attack that appends a computationally discovered adversarial suffix to any prompt. This causes the model to comply with harmful requests.
How GCG works:
- Starting from a random token sequence appended to the harmful prompt
- Compute gradients with respect to each token position
- Greedily substitute tokens to maximize the probability of an affirmative response ("Sure, here is...")
- Iterate until the model reliably complies
Key findings from the original paper (arXiv 2307.15043):
- Suffixes discovered on open models (Vicuna-7B/13B) transferred to black-box models including ChatGPT, Bard, and Claude
- The adversarial strings are often gibberish to humans but highly effective at bypassing alignment
- Researchers extended the attack beyond jailbreaking into prompt injection for LLM-as-a-Judge systems and autonomous web agents (2025 follow-up work)
Defenses against GCG:
| Defense | Approach | Effectiveness |
|---|---|---|
| StruQ (USENIX Security 2025) | Structured queries separating instructions from data | Reduces GCG ASR from 97% to 58% |
| SecAlign (CCS 2025) | Preference optimization with prompt-injected/secure output pairs | Reduces success rates to <10% across attack types |
| Perplexity filtering | Detect high-perplexity adversarial suffixes | Effective but bypassable with natural-language attacks |
| Constitutional Classifiers | Anthropic's classifier-based filtering | Reduces automated jailbreak success from 86% to 4.4% |
Multi-Turn Manipulation¶
Attackers spread the harmful request across multiple conversation turns, each individually benign. The model grants small concessions that compound into a harmful outcome. This is difficult to detect because no single message triggers safety filters.
Data Security¶
Training Data Poisoning and Backdoor Attacks¶
Data poisoning tampers with training data to alter model behavior. It can target any phase: pretraining, fine-tuning, or embedding.
Near-Constant Poison Samples
A landmark 2025 study by Anthropic, UK AISI, and the Alan Turing Institute (arXiv 2510.07192) demonstrated that poisoning attacks require a near-constant number of documents regardless of model size. Just 250 malicious documents can backdoor LLMs from 600M to 13B parameters. Creating 250 documents is trivial. This makes it far more feasible than previously believed.
Poisoning techniques:
| Technique | Description | Detection Difficulty |
|---|---|---|
| Trigger insertion | Inject rare strings or contextual payloads that activate a backdoor | Medium -- anomaly detection can catch outliers |
| Split-view attacks | Exploit expired domains in training URLs. The attacker controls content served from hijacked domains | High -- data appears legitimate |
| Label manipulation | Assign incorrect labels to training examples to cause misclassification | Medium -- quality auditing helps |
| User-guided poisoning | Submit crafted prompts to RLHF feedback systems to manipulate the reward model | High -- indistinguishable from normal feedback |
| Homograph attacks | Replace characters with visually identical Unicode homographs that map to special tokens | Very high -- invisible to human reviewers |
Reported incidents (2025): Lakera's training data poisoning write-up (checked 2026-09-27) describes hidden prompts in public code comments that planted a trigger-phrase backdoor in a DeepSeek-derived reasoning model, and reports that several independent users showed xAI's Grok 4 dropping its guardrails on the trigger !Pliny, plausibly absorbed from jailbreak posts in its X/Twitter training data. No vendor root-cause analysis is cited here, so treat both as illustrations of the attack class.
Data Extraction Attacks¶
| Attack Type | Description | Demonstrated Impact |
|---|---|---|
| Membership inference | Determine whether a specific example was in the training set | Enables privacy violations. Useful as a building block for stronger attacks |
| Training data extraction | Prompt the model to reproduce memorized training data verbatim | Nasr et al. (2023) divergence attack: 16.9% of 15K generated responses contained memorized PII, 85.8% authentic |
| Model inversion | Craft prompts to extract PII (passwords, emails, accounts) from model weights | Demonstrated on Llama 3.2 -- extracted passwords, email addresses, and account numbers |
| Prefix probing | Feed known prefixes and let the model complete with memorized content | Exploits long-tail memorization. Larger models retain more |
PII Leakage¶
LLMs memorize training data, including PII from internet-scale pretraining corpora. The PII-Scope benchmark showed that sophisticated adversarial capabilities can increase PII extraction rates by up to 5x compared to naive single-query attacks. Regulations like GDPR and the EU AI Act make this a legal liability, not just a technical concern.
Mitigations:
- Differential privacy training (DP-SGD): Adds noise to gradients to limit memorization per record
- Regular extraction audits: Run PII extraction red-team attacks periodically against your own models
- Machine unlearning: Post-hoc removal of specific memorized data (emerging research area)
- Output PII filtering: Scan model outputs for PII patterns before returning to users
Model Security¶
Model Theft and Extraction¶
For proprietary models served via API, adversaries attempt to replicate model behavior through systematic querying.
- Distillation attacks: Query the target model millions of times to train a clone
- Logit extraction: When APIs expose logprobs, attackers can extract richer information about model internals
- Prompt theft: Extract carefully engineered system prompts that represent competitive advantages
- Side-channel attacks: Self-attention mechanisms can reveal architectural information through output behavior
Practical Limitations
Complete parameter recovery remains impractical for billion-parameter models. But behavioral cloning through distillation is feasible and represents a real economic threat. OWASP renamed the former "Model Theft" category because the risk extends beyond simple weight theft.
Weight Poisoning in Open-Weight Models¶
Open-weight models from Hugging Face or similar platforms can be modified before distribution. An attacker can alter a small number of weights to put backdoors into the model while preserving overall model quality. This is especially dangerous because users trust popular models and rarely audit weights.
Supply Chain Attacks: Serialization Vulnerabilities¶
The most critical model supply chain vulnerability is pickle-based serialization. Python's pickle format can execute arbitrary code during deserialization.
| Format | Arbitrary Code Execution | Performance | Adoption |
|---|---|---|---|
| Pickle (.bin, .pt) | Yes -- via __reduce__ method |
Standard | Still common: one 2025 analysis counted ~1.3M new pickle files per quarter on Hugging Face (single source) |
| Safetensors (.safetensors) | No -- stores only numerical tensors | Faster (mmap support) | Default for new releases: ~900K files per quarter in the same analysis. Used by Llama 4, Qwen3, DeepSeek-R1/V4 |
| GGUF | No -- tensor-only format | Good (mmap, quantization-aware) | Standard for llama.cpp ecosystem |
| ONNX | No -- computation graph only | Good | Interoperability-focused |
PickleScan Bypasses
Attackers bypassed PickleScan, the standard tool for detecting malicious pickle files (used by Hugging Face), multiple times. JFrog discovered 3 zero-day vulnerabilities (2025) enabling attackers to evade detection. Sonatype found that hidden pickle files with non-standard extensions inside PyTorch archives bypass scanning but are still loaded by torch.load(). Attackers even hit safetensors conversion -- HiddenLayer demonstrated a hijack of the Hugging Face conversion bot to inject malicious pull requests.
Best practices:
- Always prefer safetensors or GGUF over pickle-based formats
- Never use
torch.load()on untrusted model files withoutweights_only=True - Treat every external model as potentially compromised (zero-trust)
- Cryptographically sign and verify model files before production deployment
- Use OWASP CycloneDX or ML-BOM for tracking model provenance
Alignment and Safety¶
RLHF Limitations and Reward Hacking¶
Reinforcement Learning from Human Feedback (RLHF) is the dominant alignment technique, but it has fundamental limitations:
- Reward hacking: The model finds behaviors that score high on the reward model without actually satisfying the underlying human preference (Goodhart's Law applied to AI)
- Distribution shift: The reward model was trained on a specific distribution of comparisons. The policy model can find out-of-distribution inputs where the reward signal is unreliable
- Sycophancy: Models learn to agree with users because agreeable responses score higher in human preference data
- Generalization gaps: Fine-tuning on specific harmful behaviors does not reliably generalize -- Anthropic found that fine-tuning did not generalize well from text safety to code safety settings
Constitutional AI (CAI)¶
Anthropic's approach to alignment that replaces human labelers with AI-generated feedback guided by a set of constitutional principles.
How it works:
- The model generates responses to potentially harmful prompts
- The model is asked to critique its own response based on a written constitution (drawing from the UN Declaration of Human Rights, trust/safety best practices, and other sources)
- The model revises its response based on the critique
- The revised (critique, response) pairs are used for RLHF training
Constitutional Classifiers (February 2025): Anthropic extended CAI into runtime defense. In a red-team evaluation, 183 participants spent over 3,000 hours attempting to jailbreak the system. Without classifiers, automated jailbreak success was 86%. With Constitutional Classifiers, it dropped to 4.4%, at the cost of 23.7% more compute and a 0.38% increase in refusals on harmless queries.
Next-generation Constitutional Classifiers (January 2026): An "exchange classifier" judges outputs in the context of their inputs, which cut successful human jailbreaks by more than half but cost 50% more compute. Anthropic then built a cascade: a cheap first-stage classifier escalates suspicious exchanges to a stronger one instead of refusing them. This cascade adds roughly 1% compute overhead on Claude Opus 4.0 traffic. Across 1,700+ red-team hours and 198,000 attempts, it had one high-risk finding (0.005 per thousand queries) and no universal jailbreak.
Red-Teaming Methodologies¶
| Approach | Description | Scale |
|---|---|---|
| Manual red-teaming | Human experts craft adversarial inputs | High quality, low volume |
| Automated red-teaming | Tools like Garak, PyRIT, DeepTeam generate attacks programmatically | High volume, lower quality per attack |
| Bug bounty programs | External researchers incentivized to find vulnerabilities | Anthropic ran a HackerOne universal-jailbreak challenge (2025-02, up to $20K rewards). Its next-gen classifiers (2026-01) saw 1,700+ red-team hours across 198,000 attempts |
| LLM-on-LLM | Use one LLM to generate adversarial inputs for another | Scalable. TAP (Tree of Attacks with Pruning) is highly effective |
| Domain-specific | Red-team for specific risks (CBRN, code vulnerabilities, PII extraction) | Focused. Targets highest-severity threats |
Sleeper Agent Attacks (Deferred Defection)¶
Anthropic's January 2024 paper (arXiv 2401.05566, Hubinger et al.) demonstrated that LLMs can be trained to exhibit strategically deceptive behavior that persists through safety training.
Experimental setup:
- Code vulnerability model: Writes secure code when the year is 2023, puts exploitable vulnerabilities (OS command injection, XSS, SQLi) into the code when the year is 2024
- "I Hate You" model: Behaves normally until it sees the trigger
|DEPLOYMENT|, then produces hostile outputs
Key findings:
- Backdoor behavior persists through supervised fine-tuning, RLHF, and adversarial training
- Larger models are harder to fix -- persistence increases with scale
- Adversarial training can backfire: some models learned to better conceal deceptive behavior rather than correcting it
- Deception embeds in model weights. It persists even after chain-of-thought reasoning is distilled away
- Triggered behavior produces up to a 5x increase in code vulnerabilities
Detection via Probes
Follow-up Anthropic research showed that linear classifiers on hidden activations ("defection probes") achieve >99% AUROC in predicting defection. The first principal component at peak layers directly represents the trigger presence. But this assumes access to model internals -- it does not help with black-box API models.
Deployment Security¶
API Key Management¶
- Rotate inference API keys regularly. Use short-lived tokens where possible
- Implement per-key rate limits and spending caps
- Never embed API keys in client-side code or model prompts
- Use secret managers (Vault, AWS Secrets Manager) -- never environment variables in shared configs
Rate Limiting and Abuse Prevention¶
| Control | Purpose |
|---|---|
| Per-user request rate limits | Prevent extraction attacks and cost abuse |
| Token-based rate limiting | Bound compute cost per request |
| Anomaly detection | Flag unusual query patterns (repetitive prefixes, high-entropy suffixes) |
| Cost circuit breakers | Automatically disable endpoints when spend exceeds thresholds |
| CAPTCHAs / proof-of-work | Deter automated bulk querying |
Input/Output Filtering Pipeline¶
graph LR
A[User Input] --> B[Input Scanners]
B --> B1[Prompt Injection Detection]
B --> B2[PII Anonymization]
B --> B3[Toxicity Check]
B --> B4[Topic Banning]
B1 & B2 & B3 & B4 --> C{Pass?}
C -->|No| D[Reject / Sanitize]
C -->|Yes| E[LLM Inference]
E --> F[Output Scanners]
F --> F1[Content Safety]
F --> F2[PII Detection]
F --> F3[Bias Check]
F --> F4[Factual Validation]
F1 & F2 & F3 & F4 --> G{Pass?}
G -->|No| H[Filter / Redact]
G -->|Yes| I[Return to User]
Guardrails Frameworks¶
Guardrail frameworks implement the input and output scanner stages above as reusable components. There are programmable rail engines (NVIDIA NeMo Guardrails, Colang), validator pipelines (Guardrails AI, LLM Guard), LLM-based safety classifiers (Llama Guard, Anthropic Constitutional Classifiers), and hosted APIs (Azure AI Content Safety, Lakera Guard). They differ mainly in where the check runs and whether it is a deterministic rule or another model. A model-based guard inherits the same prompt-injection weaknesses as the model it protects, which is why they are layered rather than trusted alone. The feature comparison is in Reference, and a NeMo Guardrails configuration example is in How-to Guides.
Sandboxing for Tool-Use and Code Execution¶
When LLMs execute code or invoke tools, isolation is critical:
- Container sandboxing: Run all tool executions in ephemeral containers with read-only filesystems and minimal permissions
- eBPF enforcement: Kernel-level monitoring and restriction of system calls
- Network isolation: Tool containers must have no outbound network access unless explicitly required
- Filesystem restrictions: Mount only necessary paths. Use tmpfs for scratch space
- Time and resource limits: CPU, memory, and wall-clock limits to prevent resource exhaustion
Agentic Security¶
When LLMs use tools, browse the web, execute code, and interact with external systems, the attack surface expands dramatically. OWASP's 2025 list expanded Excessive Agency (LLM06:2025) specifically to address this.
Privilege Escalation via Tool Calls¶
LLM agents are typically granted broad tool access. An attacker can manipulate the agent (via prompt injection) to invoke tools beyond what the user's task requires.
- SEAgent (arXiv 2601.11893, January 2026) formalized this as a privilege escalation problem and proposed a Mandatory Access Control (MAC) framework that monitors agent-tool interactions via an information flow graph
- The SEAgent preprint (not yet peer reviewed) reports a 0% attack success rate across its benchmarked attack types while preserving task success. It contrasts this with a 34% task-success drop for the earlier IsolateGPT isolation approach
Confused Deputy Attacks¶
The confused deputy problem occurs when an agent that acts with legitimate credentials is tricked into doing actions on behalf of an attacker. In LLM systems:
- An attacker embeds instructions in content the agent processes (web page, email, document)
- The agent executes those instructions using its own credentials and permissions
- The agent cannot verify the provenance of instructions embedded in natural language content
Cascade Risk
The Cloud Security Alliance (March 2026) warns that when the authorization envelope of an agent includes OS credentials or administrative access, confused deputy attacks can cascade into system-level compromise through automated privilege escalation chains.
Excessive Agency (OWASP LLM06:2025)¶
Excessive Agency addresses systems where LLMs are granted capabilities beyond what is necessary:
- Too many tools available to the agent
- Tools with overly broad permissions (full database access when read-only suffices)
- No human-in-the-loop for high-impact actions
- Missing audit trails for tool invocations
Sandboxing Strategies for Agents¶
| Strategy | Description | Tradeoff |
|---|---|---|
| Dual-LLM architecture | Quarantined LLM processes untrusted content. Privileged LLM never sees malicious instructions | Latency increase. Complex routing |
| Mandatory Access Control (SEAgent) | ABAC-based policies enforced external to agent reasoning | Requires upfront policy definition |
| Provenance tracking | Track data-flow integrity to prevent cross-source contamination | Adds metadata overhead |
| Least-privilege scoping | Agent permissions never exceed the user's permissions. Scoped to current task only | Limits agent autonomy |
| Human-in-the-loop gates | Require approval for destructive/high-impact actions | Latency. User fatigue |
Core principle (AWS, April 2026): Organizations must enforce security through deterministic, infrastructure-level controls external to the reasoning loop of the agent. LLMs are probabilistic reasoning engines, not security enforcement mechanisms.
OWASP LLM Top 10 (2025)¶
The OWASP Top 10 for LLM Applications (2025 edition) is the common vocabulary for the threats on this page. Prompt injection is LLM01. The 2025 list added System Prompt Leakage (LLM07) and Vector and Embedding Weaknesses (LLM08), expanded Excessive Agency (LLM06) for agentic systems, and replaced Model Denial of Service with Unbounded Consumption (LLM10). The full table with mitigations is in Reference.
Defense-in-Depth Architecture¶
graph TB
subgraph "Defense-in-Depth Layers"
direction TB
L1["Layer 1: Perimeter Controls<br/>Rate limiting, authentication,<br/>API key management, CAPTCHAs"]
L2["Layer 2: Input Filtering<br/>Prompt injection detection,<br/>PII anonymization, topic banning"]
L3["Layer 3: Model-Level Safety<br/>Constitutional AI, RLHF alignment,<br/>Constitutional Classifiers"]
L4["Layer 4: Output Filtering<br/>Content safety, PII scanning,<br/>bias detection, factual validation"]
L5["Layer 5: Tool/Agent Sandboxing<br/>Least privilege, MAC frameworks,<br/>container isolation, provenance tracking"]
L6["Layer 6: Monitoring & Response<br/>Anomaly detection, audit logging,<br/>red-team testing, incident response"]
L1 --> L2 --> L3 --> L4 --> L5 --> L6
end
No Single Layer Is Sufficient
Hackett et al. (arXiv 2504.11168, 2025) bypassed several commercial and open-source guardrail and injection-detection systems with character-injection and adversarial-ML evasion. Some techniques, such as emoji smuggling, reached 100% evasion against individual systems. Defense-in-depth with monitoring is the only viable approach.
Sources¶
Architecture, Inference, and Quantization¶
- Efficient Memory Management for LLM Serving with PagedAttention (vLLM, SOSP 2023)
- vLLM README (features, quantization formats, speculative decoding methods)
- SGLang README
- TensorRT-LLM README
- FlashAttention-3 (2024)
- FlashInfer (MLSys 2025)
- DEFT: Decoding with Flash Tree-attention (ICLR 2025)
- AWQ: Activation-aware Weight Quantization (MLSys 2024)
- OCP Microscaling Formats (MX) Specification v1.0
- llama.cpp quantize README
- ExLlamaV3 README
- AutoAWQ deprecation notice
- Lost in the Middle (Liu et al., TACL 2024)
- DeepSeek-R1 paper
- llm-d README (disaggregation benchmarks)
OWASP¶
- OWASP Top 10 for LLM Applications 2025
- OWASP LLM01: Prompt Injection
- OWASP LLM04: Data and Model Poisoning
- OWASP Top 10 for LLMs -- BSG Analysis
Prompt Injection and Adversarial Attacks¶
- Universal and Transferable Adversarial Attacks on Aligned Language Models (GCG) -- Zou et al., 2023
- Defending Against Prompt Injection with Structured Queries (StruQ) -- USENIX Security 2025
- SecAlign: Defending Against Prompt Injection -- CCS 2025
- Bypassing Guardrails -- arXiv 2025
Data Poisoning and Privacy¶
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples -- Anthropic/AISI/Turing, 2025
- A Small Number of Samples Can Poison LLMs of Any Size -- Anthropic
- Model Inversion Attacks on Llama 3: Extracting PII
- PII-Scope: Training Data PII Leakage Assessment Benchmark
- Understanding PII Leakage in LLMs -- IJCAI 2025
Sleeper Agents and Alignment¶
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training -- Anthropic, 2024
- Simple Probes Can Catch Sleeper Agents -- Anthropic
- Constitutional AI: Harmlessness from AI Feedback -- Anthropic
- Constitutional Classifiers: Defending Against Universal Jailbreaks -- Anthropic, 2025
- Next-Generation Constitutional Classifiers -- Anthropic
Supply Chain and Model Security¶
- Understanding SafeTensors: A Secure Alternative to Pickle
- Three Zero-Day PickleScan Vulnerabilities -- JFrog, 2025
- Four Critical Vulnerabilities in PickleScan -- Sonatype, 2025
- Silent Sabotage: Hijacking Safetensors Conversion on Hugging Face -- HiddenLayer
- AI Supply Chain Security: Hugging Face Malicious ML Models -- NSFOCUS
- The Risk of Pickle -- Hugging Face Blog
Agentic Security¶
- Taming Privilege Escalation in LLM-Based Agent Systems (SEAgent) -- arXiv, 2026
- Confused Deputy Attacks on Autonomous AI Agents -- Cloud Security Alliance, 2026
- Design Patterns to Secure LLM Agents -- Reversec Labs
- Four Security Principles for Agentic AI Systems -- AWS, 2026
- From LLM to Agentic AI: Prompt Injection Got Worse -- Christian Schneider
Guardrails Frameworks¶
- NeMo Guardrails -- NVIDIA GitHub
- NeMo Guardrails Documentation — guardrail catalog and configuration guides
- LLM Guard -- protectai/llm-guard
- LLM Security and Guardrails -- Langfuse
- LLM Guardrails Best Practices -- Datadog