LLM Fundamentals¶
Summary
Large Language Models (LLMs) are transformer neural networks trained on trillions of tokens to predict the next token. This topic covers the whole stack: transformer internals (attention, MoE, positional encoding), the training pipeline (pretraining, SFT, preference alignment, RL for reasoning), numeric and quantization formats (BF16, FP8, MXFP4/NVFP4, GGUF, AWQ/GPTQ, EXL3), inference serving (vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, MLX), fine-tuning (LoRA/QLoRA), RAG, structured output, and LLM security. As of 2026-09, 1M-token context windows are the frontier default, large open-weight models are almost all MoE, and serving has moved to prefix caching, speculative decoding, and disaggregated prefill/decode.
Summary¶
LLMs are built on the transformer architecture (Vaswani et al., 2017). Self-attention lets every token weigh its relationship to every other token in parallel. When scaled to billions of parameters and trillions of training tokens, capabilities such as reasoning, code generation, and multi-step tool use emerge. Since 2024 these capabilities have been pushed further by reinforcement learning on verifiable rewards and by spending more compute at inference time ("thinking").
The modern LLM lifecycle has four phases:
- Pretraining: next-token prediction on trillions of tokens builds broad knowledge
- Supervised Fine-Tuning (SFT): instruction following on curated prompt-response pairs
- Alignment and RL: RLHF, DPO, or RLAIF for preferences, then RL with verifiable rewards (for example GRPO) for reasoning
- Deployment: quantization, serving engines, distributed inference, and scaling
This is how those phases connect to the artifacts and runtimes covered in this topic:
flowchart LR
subgraph TRAIN["Training"]
PT["Pretraining<br/>(next-token loss, BF16/FP8)"] --> SFT["SFT"]
SFT --> AL["Alignment + RL<br/>(DPO / RLHF / GRPO)"]
end
AL --> CKPT[("Safetensors checkpoint<br/>(Hugging Face Hub)")]
CKPT --> QZ["Quantize / convert<br/>(llm-compressor, GPTQModel,<br/>llama-quantize, mlx_lm.convert)"]
CKPT --> FT["PEFT fine-tune<br/>(LoRA / QLoRA)"]
FT --> CKPT
QZ --> SRV["Serving engine<br/>(vLLM, SGLang, TensorRT-LLM,<br/>llama.cpp, Ollama, MLX)"]
CKPT --> SRV
SRV --> APP["Applications<br/>(RAG, agents, structured output,<br/>guardrails)"]
Key Facts¶
| Fact | Value (as of 2026-09-25) |
|---|---|
| Core architecture | Decoder-only transformer. Pre-norm RMSNorm, SwiGLU FFN, RoPE, GQA/MLA attention. MoE for large models. Hybrid attention/SSM layers are emerging |
| Latest Version (date): serving engines | vLLM 0.30.0 (2026-09-22). SGLang 0.5.20 (2026-09-18). TensorRT-LLM 1.2.1 stable (2026-04-20). llama.cpp b11100 (2026-09-22). Ollama v0.34.x (2026-09). mlx 0.32.2 (2026-08-25) |
| Latest Version (date): tooling | Transformers 5.17.0 (2026-09-09). PEFT 0.21.0 (2026-09-15). llm-compressor 0.14.0 (2026-09-22). NVIDIA Dynamo 1.5.0 (2026-09-19). llm-d v0.7 (2026-05) |
| Frontier context window | 1M tokens (Claude Opus 5.5 / Sonnet 5 / Fable 5.1, Gemini 3.1 Pro, DeepSeek-V4). Claude Haiku 4.5 is 200K |
| Frontier API pricing range | $1–$10 input / $5–$50 output per MTok (Claude Haiku 4.5 to Claude Fable 5.1 and GPT-6 Astra) |
| Largest open-weight models | DeepSeek-V4-Pro (1.6T total / 49B active, MIT, per DeepSeek's release note). Qwen3.5 397B-A17B (Apache-2.0) |
| Production precision | BF16 baseline. FP8 on Hopper. NVFP4/MXFP4 on Blackwell. GGUF Q4_K_M / MLX 4-bit locally |
| Licenses (engines) | vLLM, SGLang, TensorRT-LLM, Dynamo, llm-d: Apache-2.0. llama.cpp, Ollama, MLX, ExLlamaV3: MIT |
| Notable deprecations | Hugging Face TGI in maintenance mode (2025-12-11). AutoAWQ deprecated (merged into llm-compressor). AutoGPTQ unmaintained (use GPTQModel). vLLM guided_* fields removed in v0.12.0 |
Full tables are in Reference.
Key Concepts at a Glance¶
| Concept | What It Is | Details |
|---|---|---|
| Transformer | Core neural network architecture | Explanation |
| Self-Attention | Mechanism to weigh token relationships (QKV) | Explanation |
| Attention variants | MHA, GQA, MLA, sliding-window, sparse/hybrid | Explanation |
| FFN / SwiGLU | Feed-forward network with gated activation | Explanation |
| RMSNorm / Pre-LN | Normalization and its placement in transformer blocks | Explanation |
| MoE | Mixture of Experts with sparse activation | Explanation |
| Tokenization | Breaking text into sub-word units (BPE) | Explanation |
| RoPE / ALiBi | Positional encoding methods | Explanation |
| FP16 / BF16 / FP8 / FP4 | Floating-point precision formats | Explanation · table |
| Quantization | Reducing weight precision (BF16 to 8/4-bit) | Explanation |
| GGUF / GPTQ / AWQ / EXL3 | Model file and quantization formats | Explanation · GGUF table |
| MLX | Apple Silicon ML framework | Explanation |
| Training Pipeline | Pretraining, SFT, RLHF/DPO, RL for reasoning | Explanation |
| Distillation | Teacher-student model compression | Explanation |
| Scaling Laws | Training and inference-time compute scaling | Explanation |
| Model Merging | Combining fine-tuned models (TIES, DARE, SLERP) | Explanation |
| KV Cache | Key-value cache for inference speedup | Explanation |
| Speculative Decoding | Draft-and-verify decoding (EAGLE-3, MTP, DFlash 2) | Explanation · deep dive: DFlash 2 |
| VRAM Estimation | Calculating GPU memory requirements | How-to |
| GPU Selection | Hardware selection guide | How-to · specs |
| LoRA / QLoRA | Parameter-efficient fine-tuning | How-to |
| RAG | Retrieval-Augmented Generation | How-to |
| vLLM / SGLang / TensorRT-LLM | Production serving engines | How-to · versions |
| Benchmarks | MMLU-Pro, GPQA, SWE-bench, LMArena, and more | Reference |
| Structured Output | Constrained decoding, JSON schema | How-to |
| Safety / Guardrails | Content filtering, prompt injection defense | How-to · threat model |
| Glossary | TTFT, prefill, decode, PagedAttention, and more | Reference |
Evaluation¶
| Dimension | Rating | Notes |
|---|---|---|
| Maturity | High | The transformer has been battle-tested since 2017. MoE has dominated large open models since 2024–2025 |
| Ecosystem | Massive | Hugging Face, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, MLX. Engines ship day-0 support for major open models |
| Accessibility | Improving | QLoRA fine-tunes 65–70B models on a single 48–80 GB GPU. Hosted APIs start at ~$1/MTok input |
| Local Inference | Strong | GGUF + llama.cpp/Ollama and MLX run 8B–32B models on consumer hardware, and large MoE models on 128–512 GB Macs |
| Rate of change | Very high | Model lineups, prices, and engine versions change monthly. Recheck Reference before relying on numbers |
Pros
- Open, well-documented architecture with mature open-source training and serving stacks
- Strong open-weight options (DeepSeek, Qwen, gpt-oss, Llama, Gemma) under permissive licenses (MIT, Apache-2.0) or custom community licenses
- Serving efficiency keeps compounding (PagedAttention, prefix caching, FP8/FP4, speculative decoding, disaggregation)
Cons
- Memory (weights + KV cache) rather than compute is the usual bottleneck. MoE saves FLOPs, not VRAM
- Quality after aggressive quantization and long-context behavior must be validated per model and task
- Security is unsolved at the model level. Prompt injection needs defense-in-depth outside the model
Topic Map¶
- How-to Guides: choose a serving engine (with a decision flowchart), estimate VRAM, pick GPUs and quantization formats, fine-tune with PEFT, scale out with Dynamo/llm-d, build RAG, enforce JSON schemas, configure guardrails, production checklist, and CLI recipes
- Reference: model landscape and context windows, numeric formats, GGUF bits/weight, engine and tooling versions, GPU specs, benchmarks, vector databases, guardrail frameworks, OWASP LLM Top 10, and glossary
- Explanation: transformer internals, attention variants, tokenization, training pipeline, MoE, quantization internals, inference optimization (serving architecture, KV cache, speculative decoding), MLX, distillation, scaling laws, merging, post-transformer architectures, post-training math, GPU parallelism, multimodality, and the LLM security threat model
- Source notes: Ref: Alisa's Book of LLMs (derivation-heavy transformer, post-training, and parallelism notes) and Ref: Alisa's Math Notes (probability and statistics for ML)
Related Topics¶
- DFlash 2: block-diffusion speculative decoding, built on the KV cache and draft-verify ideas here
- DFlash 2 vs EAGLE-3 vs MTP: which speculative drafter to pick
- LLM Inference domain: inference-specific research
- AI Platform Engineering: platforms that host these models
- Zero Data Retention: provider data-retention controls for hosted LLM APIs
- LLM Wiki and Jev: applications that consume LLM structured output and context
- PostgreSQL: home of pgvector for RAG
Sources¶
Official Documentation and Repositories¶
- Attention Is All You Need (Vaswani et al., 2017): the original transformer paper
- Claude models overview — Anthropic
- OpenAI API pricing and Gemini API pricing
- DeepSeek V4 Preview Release — DeepSeek API Docs
- vLLM · vLLM docs · PagedAttention paper
- SGLang · TensorRT-LLM · llama.cpp · Ollama
- MLX Framework — Apple · MLX GitHub Repository · mlx-lm · Exploring LLMs with MLX on M5 — Apple ML Research
- NVIDIA Dynamo · llm-d
- Text Generation Inference (maintenance mode)
- ExLlamaV3 · AutoAWQ deprecation notice
- Quantization Overview — Hugging Face Transformers docs: GGUF/AWQ/GPTQ/bitsandbytes methods compared
- PEFT — Hugging Face GitHub
- NeMo Guardrails — NVIDIA GitHub
- OWASP Top 10 for LLM Applications 2025
Source Notes¶
- Alisa's Math Notes (Notion): probability, statistics, combinatorics, and MLE reference for ML practitioners. See ai-agents/llm-fundamentals/ref - alisa-math-notes
- Alisa's Book of LLMs (Notion): derivation-heavy transformer internals, post-training (PPO/RLHF/GRPO/DPO), parallelism, and multimodality. See ai-agents/llm-fundamentals/ref - alisa-book-of-llms
Articles and Guides¶
- Large Language Model — Wikipedia
- Transformer (deep learning architecture) — Wikipedia
- How Transformers Work — DataCamp
- What Are LLMs — IBM
- Transformer Explainer — Georgia Tech
- How Do Transformers Work — Hugging Face
- MoE LLMs — Cameron R. Wolfe
- DeepSeekMoE Paper
- MoE Infrastructure — Introl
- MoE Explained — LocalAIMaster
- MoE Powers Frontier Models — NVIDIA
- Comprehensive GGUF Analysis — Furkan Gozukara
- AI Quantization Guide 2025 — Local AI Zone
- LLM Quantization: BF16 vs FP8 vs INT4 — AIMultiple
- GGUF Q4 Q8 FP16 Guide — D-Central
- Picking the Right Size Brain — InstaSD
- LLM Quantization Explained 2026 — VRLA Tech
- Quantization Methods Compared — ai.rs
- Quantization Formats — CraftRigs
- Knowledge Distillation — IBM
- Student-Teacher Distillation Guide — DEV Community
- Knowledge Distillation for LLMs — Newline
- DPO — Cameron R. Wolfe
- RLHF Explained — IntuitionLabs
- LLM Training Methodologies 2025 — Klizos
- SFT Guide — Thunder Compute
- LoRA and QLoRA — Analytics Vidhya
- Fine-Tuning Infrastructure — Introl
- Efficient Fine-Tuning with LoRA — Databricks
- vLLM Production Deployment — Introl
- vLLM Deep Dive — martinuke0
- LLM Inference Optimization — Clarifai
- Mastering LLM Inference Optimization — NVIDIA
- KV Cache Management Survey
- Optimizing Inference — Hugging Face
- Context Window & Token Guide — QubitTool
- LLM Tokenizers — DigitalOcean
- LLM Serving Guide — Inference.net
Architecture Internals¶
- Transformer Design Guide Part 2 — Rohit Bandaru
- LLaMA Components: RMSNorm, SwiGLU, RoPE — Michael Brenndoerfer
- Advanced Transformer Architectures — Nebius Academy
- Pre-LN vs Post-LN — APXML
- Positional Embeddings: RoPE & ALiBi — Towards Data Science
- ALiBi Deep Dive: Interpolation vs Extrapolation — SambaNova
- FP16 vs BF16 Explained — Furkan Gozukara
- BF16 vs FP16 Key Differences — Bitfern
Model Merging¶
- Model Merging for LLMs — NVIDIA
- Merge LLMs with MergeKit — Hugging Face
- Model Merging Survey — Cameron R. Wolfe
VRAM & GPU¶
- Calculating GPU Memory for LLMs — BentoML
- How Much VRAM for Inference — Modal
- How Much VRAM for Fine-Tuning — Modal
RAG¶
Benchmarks¶
- LLM Benchmarks Compared — LXT
- AI Benchmarks Guide — Analytics Vidhya
- 30 LLM Evaluation Benchmarks — Evidently AI
Structured Output¶
Safety & Guardrails¶
- NeMo Guardrails — NVIDIA GitHub
- LLM Security & Guardrails — Langfuse
- Bypassing Guardrails (2025 Research) — arXiv
Questions¶
- How far will hybrid attention/SSM designs (Qwen3.5, Nemotron-H) replace full attention at frontier scale? See Explanation.
- What is the practical floor for quantization (NVFP4, 2–3 bpw EXL3/IQ quants) before quality degrades unacceptably for agentic and tool-use workflows? See Explanation and Reference.
- Disaggregated prefill/decode now ships in vLLM, SGLang, TensorRT-LLM, Dynamo, and llm-d. At what scale does it beat colocated serving with chunked prefill? See Explanation.
- How will Apple's MLX ecosystem evolve with M5 Pro/Max memory and bandwidth (up to 128 GB and 614 GB/s on M5 Max, see Reference), and will an M5 Ultra follow?