LLM Operations¶
Task-oriented guides for running LLMs: picking a serving engine, sizing VRAM and GPUs, choosing a quantization format, fine-tuning with PEFT, scaling out, building RAG, enforcing structured output, adding guardrails, preparing for production, and copy-paste CLI recipes.
Companion pages
The why behind these tasks (KV cache, speculative decoding, FlashAttention, batching, quantization internals) is in Explanation. Version numbers, GPU specs, benchmark and vector-database tables are in Reference.
Serving Engines¶
Pick an engine first: it decides your model format, quantization options, and scaling path. Current versions and licenses are in Reference.
vLLM¶
Open-source inference engine optimized for high throughput (0.30.0, 2026-09-22). Its core innovation is PagedAttention. The original vLLM paper (SOSP 2023) showed that earlier serving systems wasted 60–80% of KV cache memory to fragmentation, and vLLM cut that to under 4%.
- The 2023 launch benchmarks showed up to 24x the throughput of Hugging Face Transformers and up to 3.5x that of early TGI. Treat these as historical, since all engines have improved since
- OpenAI-compatible API (plus Anthropic Messages API and gRPC) out of the box
- Continuous batching, chunked prefill, automatic prefix caching, disaggregated prefill/decode
- Speculative decoding (n-gram, suffix, EAGLE, DFlash) and structured outputs (XGrammar, llguidance)
- Broad hardware: NVIDIA, AMD, Intel GPUs, CPUs, and plugins for TPU, Gaudi, Ascend, and others
Distributed parallelism in vLLM:
| Strategy | When to Use | Flag |
|---|---|---|
| Tensor Parallelism (TP) | Model too large for one GPU, fits one node | --tensor-parallel-size 4 |
| Pipeline Parallelism (PP) | Model too large for one node | --pipeline-parallel-size <nodes> |
| Data Parallelism (DP) | Model fits, need more throughput | --data-parallel-size N |
| Expert Parallelism (EP) | Large MoE models | --enable-expert-parallel (with TP/DP) |
Multi-node deployments typically use Ray. Single-node runs use Python multiprocessing.
TensorRT-LLM¶
NVIDIA's inference library (stable 1.2.1, 2026-04-20. 1.3.0 release candidates through 2026-09). It is now PyTorch-based, with fused kernels, CUDA graphs, FP8/NVFP4 Tensor Core paths, large-scale expert parallelism, and disaggregated serving.
trtllm-serve <hf-model-id>starts an OpenAI-compatible server directly from a Hugging Face checkpoint. The old mandatory per-model "engine build" step is no longer required for the default PyTorch workflow- Best raw performance on NVIDIA Hopper and Blackwell, especially with NVIDIA's pre-quantized FP8/NVFP4 checkpoints
- Integrates with NVIDIA Dynamo and Triton Inference Server
- Older published figures (for example >10,000 output tok/s on H100 FP8 at 64 concurrent requests) are configuration-specific. Re-benchmark on your model
SGLang¶
High-performance serving (0.5.20, 2026-09-18) with RadixAttention, a radix-tree prefix cache that reuses KV cache across requests. It also has a zero-overhead scheduler, prefill/decode disaggregation, large-scale expert parallelism, and day-0 support for many open models (DeepSeek-V4, Kimi K3, Nemotron 3). Best for:
- Agentic workflows with repeated prefixes
- RAG systems with shared context
- Multi-turn conversations
- Frontier-scale MoE deployments (DeepSeek and Kimi on GB200/GB300 NVL72)
Ollama¶
Single-command local inference built on llama.cpp, with a CLI, a REST API on port 11434, and Python/JS client libraries (v0.34.x as of 2026-09).
Best for prototyping, local development, and personal use. Recent releases add ollama launch <integration> to wire local models into coding agents. It is not designed for high-concurrency production serving.
LM Studio¶
GUI-based local inference for GGUF and MLX models on macOS, Windows, and Linux. It downloads models from Hugging Face, runs them with one click, and can expose a local OpenAI-compatible server. Its audience is similar to Ollama's, but with a visual interface.
Engine Selection Guide¶
| Scenario | Recommended Engine |
|---|---|
| Fast time-to-serve, OpenAI-compatible, broad model and hardware support | vLLM |
| Lowest latency / highest throughput on NVIDIA Hopper or Blackwell | TensorRT-LLM (optionally under Dynamo) |
| Agentic / RAG with heavy prefix sharing, frontier MoE at scale | SGLang |
| Long multi-turn conversations | vLLM or SGLang with prefix caching. (TGI has been in maintenance mode since 2025-12, so avoid it for new deployments) |
| Local prototyping | Ollama or LM Studio |
| CPU, mixed CPU+GPU, or exotic hardware | llama.cpp |
| Apple Silicon | MLX LM, Ollama, or llama.cpp (Metal) |
| Consumer NVIDIA GPU, single user, very low bpw | ExLlamaV3 + TabbyAPI |
| Multi-node, disaggregated, Kubernetes | vLLM/SGLang/TensorRT-LLM under NVIDIA Dynamo or llm-d |
The same choice as a decision flowchart:
flowchart TD
Q1{"Where does it run?"} -->|"Laptop / desktop"| Q2{"Apple Silicon?"}
Q2 -->|"Yes"| MLX["MLX LM or Ollama"]
Q2 -->|"No"| Q3{"Need a GUI?"}
Q3 -->|"Yes"| LMS["LM Studio"]
Q3 -->|"No"| OLL["Ollama / llama.cpp (GGUF)"]
Q1 -->|"Server GPUs"| Q4{"More than one node or<br/>separate prefill/decode pools?"}
Q4 -->|"Yes"| ORCH["Dynamo or llm-d<br/>over vLLM / SGLang / TensorRT-LLM"]
Q4 -->|"No"| Q5{"Heavy shared prefixes<br/>(agents, RAG)?"}
Q5 -->|"Yes"| SGL["SGLang"]
Q5 -->|"No"| Q6{"NVIDIA-only and chasing<br/>peak tok/s?"}
Q6 -->|"Yes"| TRT["TensorRT-LLM"]
Q6 -->|"No"| VLLM["vLLM"]
VRAM Estimation¶
Core Formula¶
$$ \text{VRAM}_{\text{total}} = \text{Weights} + \text{KV Cache} + \text{Activations} + \text{Overhead} $$
1. Model Weights¶
$$ \text{Weight Memory} = \text{Parameters} \times \text{Bytes per Parameter} $$
| Precision | Bytes/Param | 7B Model | 13B Model | 70B Model |
|---|---|---|---|---|
| FP32 | 4 | 28 GB | 52 GB | 280 GB |
| FP16 / BF16 | 2 | 14 GB | 26 GB | 140 GB |
| FP8 / INT8 | 1 | 7 GB | 13 GB | 70 GB |
| INT4 / FP4 (Q4, AWQ, NVFP4) | ~0.5–0.6 | 3.5–4.3 GB | 6.5–8 GB | 35–43 GB |
Quick rule of thumb: ~2 GB per 1B parameters at BF16, ~0.6 GB per 1B at 4-bit including scales. For example, llama.cpp's Llama 3.1 70B Q4_K_M file is 43.1 GB. Bytes for other formats are in Reference.
2. KV Cache¶
The KV cache is the hidden memory monster. It scales linearly with sequence length, batch size, and number of layers:
$$ \text{KV Cache} = 2 \times n_{\text{layers}} \times n_{\text{kv_heads}} \times d_{\text{head}} \times \text{seq_len} \times \text{batch} \times \text{bytes} $$
For Llama 3 70B with GQA (80 layers, 8 KV heads, head dim 128), this is ~0.31 MB per token at BF16. The same model with full MHA (64 KV heads) would need ~2.6 MB per token (8x more).
KV Cache Gotcha
A model that fits comfortably at 2K context can OOM at 32K. Each 1,000 tokens adds ~0.13 GB for Llama 3 8B (GQA, 8 KV heads) but ~0.5 GB for Llama 2 7B (MHA, 32 KV heads). For 70B-class models with long context and many concurrent users, the KV cache can exceed the weight memory. FP8 KV cache halves it.
3. Activations and Overhead¶
- Activations: intermediate tensors during the forward pass. Typically 5–20% of weight memory for inference
- Framework overhead: CUDA context, memory allocator, CUDA graphs, driver. From 500 MB to 2 GB or more
- Serving engines pre-allocate a fraction of VRAM up front (vLLM
--gpu-memory-utilization, default 0.9) and fill the rest with KV cache blocks
Practical formula for inference:
$$ \text{VRAM}_{\text{inference}} \approx \text{Weight Memory} \times 1.2 + \text{KV Cache} $$
4. Training VRAM¶
Full fine-tuning with mixed precision and Adam needs roughly 16 bytes per parameter before activations:
| Component | Memory (BF16 mixed-precision training) |
|---|---|
| Model weights (BF16) | 2 bytes/param |
| Gradients (BF16) | 2 bytes/param |
| FP32 master weights | 4 bytes/param |
| Optimizer states (Adam, two FP32 moments) | 8 bytes/param |
| Activations | Variable (depends on batch size, sequence length, checkpointing) |
| Total | ~16 GB per 1B params + activations (rule of thumb) |
ZeRO/FSDP sharding divides the non-activation part across GPUs (see Explanation). QLoRA reduces the requirement to roughly 0.7–1 GB per 1B params by quantizing the frozen base to 4-bit and training only LoRA adapters.
MoE Memory¶
All experts must reside in memory even though only the top-k fire per token. DeepSeek-V3 (671B total, 37B active) needs ~700 GB at FP8 for weights alone, so it is served multi-GPU (for example 8x H200) or with expert offload. MoE saves compute per token, not memory.
GPU Hardware Selection Guide¶
Pick hardware by first computing VRAM (above), then matching the workload. Full specs for NVIDIA, AMD, consumer, and Apple Silicon parts are in Reference.
Apple Silicon's advantage is unified memory: most of system RAM is usable for weights. A 512 GB M3 Ultra Mac Studio can load models that need several 80 GB GPUs, but at ~819 GB/s instead of 3–8 TB/s, so tokens per second are much lower.
Which GPU for Which Task?¶
| Task | Recommended | Why |
|---|---|---|
| Local chat (personal use) | RTX 4090 / 5090, M4 Pro/Max, M5-family Macs | 24–32 GB VRAM (or 36–128 GB unified) handles 8B–32B at 4-bit |
| Local large MoE / 70B-class | M4 Max 128 GB, M3 Ultra, RTX PRO 6000 (96 GB) | Capacity matters more than bandwidth for fitting weights |
| Fine-tuning (QLoRA, 7B–13B) | RTX 3090/4090/5090 (24–32 GB) | Enough VRAM for QLoRA |
| Fine-tuning (QLoRA, 70B) | 1x H100/H200 or 2x A100 80GB | 70B QLoRA needs ~48–80 GB depending on sequence length |
| Production serving (<14B) | L4, L40S, or A10G | Cost-effective for smaller models |
| Production serving (70B dense) | 2–4x H100/H200 with TP, or 1x H200 at FP8 | Tensor parallelism across GPUs |
| Frontier MoE serving | 8x H200/B200 nodes, GB200/GB300 NVL72, MI300X/MI355X | Expert parallelism over fast interconnect |
| Research / large-scale training | H100/H200/B200/B300 clusters | Maximum throughput |
Choose a Quantization Format¶
Match the format to your runtime and hardware. Internals are in Explanation, and GGUF bits/weight are in Reference.
Quick Format Decision Guide¶
| Scenario | Best Format |
|---|---|
| CPU / laptop / Apple Silicon | GGUF (Q4_K_M or Q5_K_M), or MLX 4-bit on Macs |
| NVIDIA GPU, max serving throughput (Ampere/Ada) | AWQ or GPTQ INT4 with Marlin kernels, or INT8 W8A8 |
| NVIDIA H100/H200 production | FP8 (weights + activations, FP8 KV cache) |
| NVIDIA Blackwell production | NVFP4 (or MXFP4 where the model ships it, for example gpt-oss) |
| NVIDIA consumer GPU, single-user interactive | EXL3 at 3–6 bpw (ExLlamaV3) |
| Fine-tuning | bitsandbytes NF4 (QLoRA) |
| Limited VRAM (≤8 GB) | GGUF Q4_K_M (or IQ4_XS) with partial CPU offload |
| General starting point | Ollama with its default Q4_K_M tag |
Always compare the quantized model against the BF16 baseline on your own task set, especially for tool calling, JSON output, and code.
Parameter-Efficient Fine-Tuning (PEFT)¶
Full fine-tuning updates all parameters and is prohibitively expensive for large models. PEFT methods train <1% of parameters and typically retain most of full fine-tuning quality.
LoRA (Low-Rank Adaptation)¶
LoRA injects trainable low-rank matrices into each transformer layer while freezing the original weights.
How it works:
For a weight matrix $W \in \mathbb{R}^{d \times k}$, instead of updating $W$ directly, LoRA adds:
$$ W' = W + \Delta W = W + BA $$
Where $B \in \mathbb{R}^{d \times r}$ and $A \in \mathbb{R}^{r \times k}$, with rank $r \ll \min(d, k)$ (typically 8–64).
| Property | Value |
|---|---|
| Trainable params | ~0.1–1% of total (0.2% is typical for r=16 on attention projections) |
| Adapter size | Few MB to a few hundred MB (vs GB for the full model) |
| Inference cost | Zero if merged into base weights. Small if served unmerged (multi-LoRA) |
| Task switching | Swap adapter files without reloading the base model. vLLM and SGLang serve many adapters on one base |
| Quality | Competitive with full fine-tuning for most narrow tasks |
QLoRA (Quantized LoRA)¶
QLoRA loads the base model in 4-bit NormalFloat (NF4) quantization while training LoRA adapters in higher precision:
- 75–80% memory reduction vs 16-bit LoRA
- The original paper fine-tuned 65B models on a single 48 GB GPU
- Quality on par with 16-bit fine-tuning in many cases
Key innovations:
- NF4 data type: optimized for normally distributed weights
- Double quantization: compresses the quantization constants themselves
- Paged optimizers: use unified memory to page optimizer state to CPU on memory spikes
Adapter Modules¶
Small feed-forward networks inserted after attention or FFN sublayers. The base model is frozen, and only adapter weights train.
- More modular than LoRA (can mix and match per task)
- Slight inference overhead (adapters do not merge into base weights)
- Useful for multi-task serving with a shared base
Other methods (DoRA, prefix tuning, prompt tuning, IA3) are compared in Reference.
When to Use Which¶
| Method | Best For | Hardware |
|---|---|---|
| Full Fine-Tuning | Maximum quality, small models (<7B) | 8x A100/H100 or equivalent |
| LoRA | General fine-tuning, easy deployment | 1–2x A100/H100 |
| QLoRA | Large models (13B–70B), limited VRAM | Single 24–80 GB GPU |
| Adapters | Multi-task serving, modular systems | Similar to LoRA |
Tooling¶
- Hugging Face PEFT (0.21.0): canonical library.
get_peft_model()wraps any Transformers model - TRL:
SFTTrainer,DPOTrainer, andGRPOTrainerfor SFT and preference/RL fine-tuning - bitsandbytes (0.50.2): 4-bit/8-bit quantization for QLoRA
- Unsloth: custom Triton kernels. Vendor-reported ~2x faster LoRA/QLoRA training with less memory
- Axolotl: config-driven fine-tuning framework wrapping multiple methods
- MLX LM:
mlx_lm.lorafor LoRA/QLoRA on Apple Silicon
Fine-Tuning Data Preparation¶
| Aspect | Recommendation |
|---|---|
| Format | JSONL with instruction/input/output (Alpaca) or a messages array (chat format). Apply the model's own chat template |
| Quality over quantity | 1,000 high-quality examples often outperform 100,000 noisy ones |
| Diversity | Cover the range of expected inputs. Do not over-represent any pattern |
| Decontamination | Remove examples overlapping with evaluation benchmarks |
| Minimum size | LoRA/QLoRA: 500–5K for task-specific. 10K–100K for general instruction tuning |
| Validation split | Hold out 5–10%. Monitor loss for overfitting |
Fine-Tuning Pipeline¶
This is the end-to-end flow from raw examples to a deployed adapter:
graph LR
A["Collect / generate data"] --> B["Format to JSONL<br/>(chat template)"]
B --> C["Decontaminate and deduplicate"]
C --> D["Train / val split"]
D --> E{"Method?"}
E -->|"fits in VRAM"| E1["LoRA (PEFT)"]
E -->|"tight VRAM"| E2["QLoRA (NF4 base)"]
E -->|"small model, max quality"| E3["Full fine-tune"]
E1 --> F["Train with early stopping"]
E2 --> F
E3 --> F
F --> G["Evaluate: held-out set + task evals"]
G --> H["Merge adapter or serve via multi-LoRA"]
H --> I["Deploy on vLLM / SGLang / Ollama"]
Distributed Inference and Scaling¶
Parallelism Strategies¶
| Strategy | What It Splits | When to Use |
|---|---|---|
| Tensor Parallelism (TP) | Individual layer weights across GPUs | Model too large for one GPU. Keep within one NVLink domain |
| Pipeline Parallelism (PP) | Sequential layers across GPUs or nodes | Multi-node deployment |
| Data Parallelism (DP) | Replicas of the full model | High throughput, model fits one GPU or node |
| Expert Parallelism (EP) | MoE experts across GPUs | MoE models with many experts. "Wide EP" spans dozens of GPUs |
How the collectives behind these work is in Explanation.
NVIDIA Dynamo¶
Announced at GTC 2025 and open source (Apache-2.0, 1.5.0 as of 2026-09-19). Dynamo is a datacenter-scale orchestration layer above SGLang, TensorRT-LLM, and vLLM. It does not replace them.
- Disaggregated prefill and decode: separate GPU pools optimized for each phase
- KV-aware routing: sends requests to workers that already hold the matching prefix
- Multi-tier KV cache (GPU, CPU, SSD), SLA-driven autoscaling, fast cold starts
- Not needed if you run a single model on a single GPU
llm-d (Kubernetes-Native)¶
Launched in May 2025 by Red Hat, Google Cloud, IBM Research, NVIDIA, and CoreWeave. It joined the CNCF as a Sandbox project on 2026-03-24, and v0.7 shipped in 2026-05.
- Kubernetes-native distributed serving over vLLM (and SGLang)
- Prefix-cache- and load-aware routing through the Gateway API Inference Extension
- Prefill/decode disaggregation, wide expert parallelism, tiered KV offload to CPU or disk
- SLO-aware autoscaling and batch (offline) serving paths
- Project-reported results: 3x output throughput with prefix-aware routing versus round-robin (Llama 3.1 70B on MI300X), and up to 70% more tokens/sec with P/D disaggregation (gpt-oss on B200)
Multi-Model Routing¶
For production deployments with multiple models:
| Tool | Purpose |
|---|---|
| LiteLLM | Unified API gateway for 100+ LLM providers. Fallback routing, budgets |
| Envoy AI Gateway | Proxy-level routing, rate limiting, auth |
| OpenRouter | Third-party multi-model API with cost optimization |
Retrieval-Augmented Generation (RAG)¶
RAG lets an LLM access external knowledge at inference time. This reduces hallucination and enables domain-specific responses without retraining.
Architecture¶
The indexing path (bottom) runs offline. The query path (top) runs per request:
graph LR
A["User query"] --> B["Embedding model"]
B --> C["Vector search<br/>(+ BM25 hybrid)"]
D["Document corpus"] --> E["Parse + chunk"]
E --> F["Embedding model"]
F --> G[("Vector database")]
C --> G
G --> H["Top-K chunks"]
H --> R["Cross-encoder reranker"]
R --> I["Augmented prompt"]
A --> I
I --> J["LLM"]
J --> K["Grounded response + citations"]
Pipeline Steps¶
| Step | What Happens | Key Decisions |
|---|---|---|
| 1. Document Preparation | Clean, parse, and normalize source documents | Format handling (PDF, HTML, markdown), table and image extraction |
| 2. Chunking | Split documents into retrieval units | Chunk size (256–1024 tokens), overlap, strategy |
| 3. Embedding | Convert chunks to dense vectors | Model choice (for example OpenAI text-embedding-3-large, Voyage, or open models such as BGE-M3 and Qwen3-Embedding) |
| 4. Indexing | Store vectors in a vector database | Database choice (see Reference) |
| 5. Retrieval | Find the top-K chunks similar to the query | Similarity metric (cosine), hybrid search, K value, metadata filters |
| 6. Reranking | Re-score retrieved chunks for relevance | Cross-encoder reranker (Cohere Rerank, BGE reranker) or late interaction (ColBERT) |
| 7. Augmentation | Inject chunks into the LLM prompt | Prompt template design, chunk ordering (put the best evidence first or last, see "Lost in the Middle") |
| 8. Generation | LLM produces a grounded answer | Citation generation, faithfulness checking |
Chunking Strategies¶
| Strategy | Description | Best For |
|---|---|---|
| Fixed-size | Split at N tokens with overlap | Simple, predictable |
| Recursive | Split by paragraph, then sentence, then word | General-purpose (LangChain default) |
| Semantic | Split where embedding similarity between adjacent segments drops | Better coherence, at the cost of an embedding pass during chunking |
| Document-structure | Split by headings, sections, markdown structure | Structured documents (docs, wikis) |
| Proposition-based | Extract atomic facts as individual chunks | Highest precision, expensive to compute |
Hybrid Search and Reranking¶
Combining dense (semantic) and sparse (lexical/BM25) retrieval improves results. Dense search handles paraphrases and meaning, while sparse search catches exact terms, acronyms, and domain jargon. Fuse the two ranked lists with Reciprocal Rank Fusion (RRF) or weighted scores.
After initial retrieval, a cross-encoder reranker re-scores each (query, chunk) pair. Vendor and practitioner reports cite 10–30% precision gains at a cost of 50–100 ms extra latency. Measure on your own queries.
RAG vs Fine-Tuning¶
| Dimension | RAG | Fine-Tuning |
|---|---|---|
| Knowledge update | Instant (swap documents) | Requires retraining |
| Cost | Low (no training) | High (GPU hours) |
| Hallucination | Reduced (grounded in sources) | Can still hallucinate |
| Latency | Higher (retrieval + generation) | Lower (single forward pass) |
| Best for | Factual Q&A, documents, knowledge bases | Style/format changes, task specialization |
Use RAG for knowledge and fine-tuning for behavior. Delivering domain knowledge through retrieval instead of weight updates avoids most retraining cost. With 1M-token context windows, "just put the documents in the prompt" is also viable for small, stable corpora when combined with prompt caching.
Structured Output and Constrained Decoding¶
These methods make sure that LLMs produce valid JSON, function calls, or other structured formats.
Approaches¶
| Method | Guarantee | How It Works |
|---|---|---|
| Prompt engineering | Best effort | Describe the desired format in the prompt |
| JSON mode (API) | Valid JSON, not necessarily your schema | Provider constrains output to syntactically valid JSON |
| Structured outputs / strict schemas (API) | Schema-conformant | Provider compiles your JSON Schema into a decoding constraint |
| Constrained decoding (self-hosted) | 100% grammar-conformant | Mask invalid tokens at each step using a grammar or schema |
| Fine-tuning | High but not guaranteed | Train on structured input-output pairs |
The token-masking mechanism is described in Explanation. Library comparison: Reference.
Enforce a JSON Schema with vLLM¶
vLLM enables structured outputs by default and picks the backend (auto: XGrammar or guidance). Send an OpenAI-style response_format:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-8B",
"messages": [{"role": "user", "content": "Extract the city and country from: I live in Lyon."}],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "location",
"schema": {
"type": "object",
"properties": {"city": {"type": "string"}, "country": {"type": "string"}},
"required": ["city", "country"]
}
}
}
}'
For regex, choice, grammar, or structural_tag constraints, pass "structured_outputs": {...} as an extra request field. structural_tag constrains only tagged regions (for example tool calls) while the rest stays free-form, which suits agentic workflows.
vLLM API change
The legacy guided_json, guided_regex, guided_choice, guided_grammar, and guided_decoding_backend request fields were removed in vLLM v0.12.0. Migrate to structured_outputs or response_format.
Safety, Guardrails, and Content Filtering¶
Frameworks are compared in Reference, and the threat model is in Explanation.
Add Rails in Layers¶
- Input rails (before the LLM): classify prompts (Llama Guard, Prompt Shields, Lakera), mask PII, detect injection patterns.
- Retrieval rails (RAG): treat retrieved chunks and tool outputs as untrusted data. Strip or flag instructions embedded in them.
- Dialog rails (multi-turn): enforce allowed topics and flows. Watch for multi-turn escalation.
- Output rails (after the LLM): toxicity and PII filters, schema validation, and escaping before output reaches a browser, shell, or SQL.
- Outside the model: rate limits, spending caps, least-privilege tool credentials, and human approval for high-impact actions.
Configure NeMo Guardrails Injection Detection¶
NeMo Guardrails (0.24.1, 2026-09) configures rails in config.yml. This example rejects responses that contain code, SQL, template, or XSS injection payloads:
rails:
config:
injection_detection:
injections:
- code
- sqli
- template
- xss
action: reject
input:
flows:
- protect prompt
output:
flows:
- protect response
- injection detection
Prompt Injection (2025)¶
Prompt injection is the top LLM security concern (OWASP LLM01). A 2025 study (Hackett et al., arXiv 2504.11168) evaded several commercial and open-source guardrail and injection detectors with character-injection and adversarial-ML techniques:
- Some techniques, such as emoji smuggling, reached 100% evasion against individual detectors
- Multi-layered defense is the only viable approach
Defense in Depth
No single guardrail is sufficient. Combine: (1) input classification (Llama Guard or similar), (2) output filtering (toxicity, PII), (3) rate limiting + anomaly detection, (4) deterministic authorization for tools outside the model, and (5) human review for high-stakes decisions.
Production Best Practices¶
Deployment Lifecycle¶
This is the path from a local prototype to a monitored production service:
graph LR
A["Prototype with Ollama / LM Studio"] --> B["Validate with vLLM / SGLang"]
B --> C["Optimize: quantization + batching + prefix cache"]
C --> D["Load test: TTFT, ITL, throughput at target concurrency"]
D --> E["Deploy: Kubernetes + autoscaling (llm-d / Dynamo)"]
E --> F["Monitor: latency SLOs + quality evals"]
F -->|"regressions"| C
Checklist¶
Pre-Production Checklist
Model Selection
- Benchmark candidate models on your actual task distribution
- Test quantized variants (Q4_K_M, AWQ, FP8, NVFP4) against a BF16 baseline
- Validate edge cases: long inputs, multilingual, structured output, tool calls
Infrastructure
- Right-size GPU selection (H100/H200/B200 for throughput, L4/L40S for cost, Apple Silicon for private local use)
- Configure tensor parallelism if the model exceeds single-GPU VRAM
- Configure continuous batching with an appropriate max batch size / max tokens per step
- Enable prefix caching for repetitive prompt patterns
Reliability
- Deploy multiple replicas behind a load balancer (prefix-aware if possible)
- Configure autoscaling on queue depth or KV cache utilization, not just CPU/GPU utilization
- Set request timeouts and max token limits
- Implement circuit breakers and fallback to smaller or cached models
- Test failover by terminating instances under load
Monitoring
- Track time-to-first-token (TTFT), inter-token latency, tokens/second, and end-to-end latency at p50/p95/p99
- Monitor GPU utilization, VRAM usage, and KV cache occupancy and preemptions
- Log prompt/response lengths for capacity planning
- Configure alerts for latency SLO violations and OOM events
Quality
- Implement output validation (JSON schema, safety filters)
- Run periodic eval benchmarks against held-out test sets
- Monitor for model drift after updates or quantization changes
Cost Optimization¶
The savings ranges below are typical published ranges, not guarantees. Measure on your workload.
| Technique | Savings | Tradeoff |
|---|---|---|
| Quantization (BF16 to 4-bit) | 70–75% weight memory. Faster decode | Minor quality loss. Validate on your tasks |
| Speculative decoding | 1.5–3x lower latency at low batch | Drafter complexity. Gains shrink at high concurrency |
| Prefix caching | 30–60% less prefill compute for repetitive prompts | Memory for cache storage |
| Continuous batching | 3–10x throughput vs static batching | Slightly higher per-request latency |
| Spot / preemptible instances | 60–80% compute cost | Requires graceful interruption handling |
| Model distillation | 5–10x cheaper inference | Upfront distillation cost |
| Hosted API prompt caching and batch APIs | Cached input is often billed at 10% of base. Batch at ~50% | Cache TTLs and async latency |
Common Pitfalls¶
Avoid These
- Over-quantizing for your use case: Q2/Q3 works for casual chat but breaks agentic workflows, JSON output, and code generation
- Ignoring TTFT: users perceive time-to-first-token as "speed" more than tokens/second
- Static batching in production: wastes GPU cycles waiting for the longest request in the batch
- No fallback strategy: a single model endpoint is a single point of failure
- Benchmarking with synthetic data: real traffic patterns (variable lengths, bursty arrivals) behave very differently from uniform benchmarks
- Skipping load testing: KV cache OOM or preemption storms under concurrent load are the most common production failures
CLI Recipes¶
Commands were checked against the projects' current READMEs and docs (2026-09-25). Model IDs are examples. Gated models such as Llama need an accepted license and HF_TOKEN.
Ollama (Local Inference)¶
# Install (Linux/macOS script. Windows and macOS also have installers)
curl -fsSL https://ollama.com/install.sh | sh
# Pull and run a model
ollama pull gemma4
ollama run gemma4
# List local models
ollama list
# Serve the API (the desktop app starts it automatically). Default: http://localhost:11434
ollama serve
curl http://localhost:11434/api/chat -d '{
"model": "gemma4",
"messages": [{"role": "user", "content": "Why is the sky blue?"}],
"stream": false
}'
vLLM (Production Serving)¶
# Install (uv recommended by the project)
uv pip install vllm
# Serve a model across 2 GPUs with tensor parallelism
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--tensor-parallel-size 2 \
--max-model-len 8192 \
--port 8000
# Pre-quantized checkpoints are detected from their config (no flag needed)
vllm serve Qwen/Qwen3-8B-AWQ
# OpenAI-compatible endpoint
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"meta-llama/Llama-3.1-8B-Instruct","messages":[{"role":"user","content":"Hello"}]}'
TensorRT-LLM (NVIDIA)¶
# Inside the NGC TensorRT-LLM release container
trtllm-serve "nvidia/Qwen3-8B-FP8" # FP8 checkpoint needs Hopper/Ada or newer
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"nvidia/Qwen3-8B-FP8","messages":[{"role":"user","content":"Hello"}],"max_tokens":32}'
MLX LM (Apple Silicon)¶
# Install
pip install mlx-lm
# Download and convert to 4-bit (default output dir: mlx_model)
mlx_lm.convert --hf-path mistralai/Mistral-7B-Instruct-v0.3 \
-q --q-bits 4 --mlx-path ./mistral-7b-4bit
# Other quantization modes: --q-mode mxfp4 | nvfp4 | mxfp8
# Generate
mlx_lm.generate --model ./mistral-7b-4bit \
--prompt "Explain transformers" --max-tokens 500
# Fine-tune with LoRA (--data is a directory containing train.jsonl, optional valid.jsonl)
mlx_lm.lora --model ./mistral-7b-4bit \
--train --data ./data --iters 1000
llama.cpp (GGUF Inference)¶
# Build (the repo moved to the ggml-org organization)
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build && cmake --build build --config Release -j 8
# Download a GGUF from Hugging Face and chat (defaults to the Q4_K_M file)
./build/bin/llama-cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# One-shot, non-interactive generation from a local file
./build/bin/llama-cli -m ./models/model-Q4_K_M.gguf \
--single-turn -p "What is quantization?" -n 256
# OpenAI-compatible API server (-ngl: layers offloaded to GPU; default auto)
./build/bin/llama-server -m ./models/model-Q4_K_M.gguf \
--host 0.0.0.0 --port 8080 -ngl all
Quantization with llama.cpp¶
pip install -r requirements.txt
# Convert a Hugging Face model directory to a BF16 GGUF
python convert_hf_to_gguf.py ./models/my-model/ --outtype bf16 --outfile my-model-bf16.gguf
# Optional: compute an importance matrix from calibration text (improves low-bit quants)
./build/bin/llama-imatrix -m my-model-bf16.gguf -f calibration.txt -o imatrix.gguf
# Quantize
./build/bin/llama-quantize --imatrix imatrix.gguf my-model-bf16.gguf my-model-Q4_K_M.gguf Q4_K_M
Fine-Tuning with QLoRA (Hugging Face)¶
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model
# 4-bit NF4 quantization config
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16,
)
model_id = "meta-llama/Llama-3.1-8B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, quantization_config=bnb_config)
lora_config = LoraConfig(
r=16, # rank
lora_alpha=32, # scaling
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
lora_dropout=0.05,
task_type="CAUSAL_LM",
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters() # roughly 0.2% trainable for this config
# Train with trl.SFTTrainer(model=model, train_dataset=..., ...)
Sources¶
- vLLM README, structured outputs docs, and PagedAttention paper
- SGLang README
- TensorRT-LLM quick start (
trtllm-serve) - Ollama README
- llama.cpp README, cli, and quantize docs
- mlx-lm README and LoRA guide
- NVIDIA Dynamo and llm-d
- Text Generation Inference maintenance-mode notice
- QLoRA paper (Dettmers et al., 2023) and PEFT
- NeMo Guardrails
- Bypassing Prompt Injection and Jailbreak Detection in LLM Guardrails (arXiv 2504.11168)
- Calculating GPU Memory for LLMs — BentoML
- vLLM Production Deployment — Introl (source of the Stripe case study)