Skip to content

LLM Fundamentals

Summary

Large Language Models (LLMs) are transformer neural networks trained on trillions of tokens to predict the next token. This topic covers the whole stack: transformer internals (attention, MoE, positional encoding), the training pipeline (pretraining, SFT, preference alignment, RL for reasoning), numeric and quantization formats (BF16, FP8, MXFP4/NVFP4, GGUF, AWQ/GPTQ, EXL3), inference serving (vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, MLX), fine-tuning (LoRA/QLoRA), RAG, structured output, and LLM security. As of 2026-09, 1M-token context windows are the frontier default, large open-weight models are almost all MoE, and serving has moved to prefix caching, speculative decoding, and disaggregated prefill/decode.

Summary

LLMs are built on the transformer architecture (Vaswani et al., 2017). Self-attention lets every token weigh its relationship to every other token in parallel. When scaled to billions of parameters and trillions of training tokens, capabilities such as reasoning, code generation, and multi-step tool use emerge. Since 2024 these capabilities have been pushed further by reinforcement learning on verifiable rewards and by spending more compute at inference time ("thinking").

The modern LLM lifecycle has four phases:

  1. Pretraining: next-token prediction on trillions of tokens builds broad knowledge
  2. Supervised Fine-Tuning (SFT): instruction following on curated prompt-response pairs
  3. Alignment and RL: RLHF, DPO, or RLAIF for preferences, then RL with verifiable rewards (for example GRPO) for reasoning
  4. Deployment: quantization, serving engines, distributed inference, and scaling

This is how those phases connect to the artifacts and runtimes covered in this topic:

flowchart LR
    subgraph TRAIN["Training"]
        PT["Pretraining<br/>(next-token loss, BF16/FP8)"] --> SFT["SFT"]
        SFT --> AL["Alignment + RL<br/>(DPO / RLHF / GRPO)"]
    end
    AL --> CKPT[("Safetensors checkpoint<br/>(Hugging Face Hub)")]
    CKPT --> QZ["Quantize / convert<br/>(llm-compressor, GPTQModel,<br/>llama-quantize, mlx_lm.convert)"]
    CKPT --> FT["PEFT fine-tune<br/>(LoRA / QLoRA)"]
    FT --> CKPT
    QZ --> SRV["Serving engine<br/>(vLLM, SGLang, TensorRT-LLM,<br/>llama.cpp, Ollama, MLX)"]
    CKPT --> SRV
    SRV --> APP["Applications<br/>(RAG, agents, structured output,<br/>guardrails)"]

Key Facts

Fact Value (as of 2026-09-25)
Core architecture Decoder-only transformer. Pre-norm RMSNorm, SwiGLU FFN, RoPE, GQA/MLA attention. MoE for large models. Hybrid attention/SSM layers are emerging
Latest Version (date): serving engines vLLM 0.30.0 (2026-09-22). SGLang 0.5.20 (2026-09-18). TensorRT-LLM 1.2.1 stable (2026-04-20). llama.cpp b11100 (2026-09-22). Ollama v0.34.x (2026-09). mlx 0.32.2 (2026-08-25)
Latest Version (date): tooling Transformers 5.17.0 (2026-09-09). PEFT 0.21.0 (2026-09-15). llm-compressor 0.14.0 (2026-09-22). NVIDIA Dynamo 1.5.0 (2026-09-19). llm-d v0.7 (2026-05)
Frontier context window 1M tokens (Claude Opus 5.5 / Sonnet 5 / Fable 5.1, Gemini 3.1 Pro, DeepSeek-V4). Claude Haiku 4.5 is 200K
Frontier API pricing range $1–$10 input / $5–$50 output per MTok (Claude Haiku 4.5 to Claude Fable 5.1 and GPT-6 Astra)
Largest open-weight models DeepSeek-V4-Pro (1.6T total / 49B active, MIT, per DeepSeek's release note). Qwen3.5 397B-A17B (Apache-2.0)
Production precision BF16 baseline. FP8 on Hopper. NVFP4/MXFP4 on Blackwell. GGUF Q4_K_M / MLX 4-bit locally
Licenses (engines) vLLM, SGLang, TensorRT-LLM, Dynamo, llm-d: Apache-2.0. llama.cpp, Ollama, MLX, ExLlamaV3: MIT
Notable deprecations Hugging Face TGI in maintenance mode (2025-12-11). AutoAWQ deprecated (merged into llm-compressor). AutoGPTQ unmaintained (use GPTQModel). vLLM guided_* fields removed in v0.12.0

Full tables are in Reference.

Key Concepts at a Glance

Concept What It Is Details
Transformer Core neural network architecture Explanation
Self-Attention Mechanism to weigh token relationships (QKV) Explanation
Attention variants MHA, GQA, MLA, sliding-window, sparse/hybrid Explanation
FFN / SwiGLU Feed-forward network with gated activation Explanation
RMSNorm / Pre-LN Normalization and its placement in transformer blocks Explanation
MoE Mixture of Experts with sparse activation Explanation
Tokenization Breaking text into sub-word units (BPE) Explanation
RoPE / ALiBi Positional encoding methods Explanation
FP16 / BF16 / FP8 / FP4 Floating-point precision formats Explanation · table
Quantization Reducing weight precision (BF16 to 8/4-bit) Explanation
GGUF / GPTQ / AWQ / EXL3 Model file and quantization formats Explanation · GGUF table
MLX Apple Silicon ML framework Explanation
Training Pipeline Pretraining, SFT, RLHF/DPO, RL for reasoning Explanation
Distillation Teacher-student model compression Explanation
Scaling Laws Training and inference-time compute scaling Explanation
Model Merging Combining fine-tuned models (TIES, DARE, SLERP) Explanation
KV Cache Key-value cache for inference speedup Explanation
Speculative Decoding Draft-and-verify decoding (EAGLE-3, MTP, DFlash 2) Explanation · deep dive: DFlash 2
VRAM Estimation Calculating GPU memory requirements How-to
GPU Selection Hardware selection guide How-to · specs
LoRA / QLoRA Parameter-efficient fine-tuning How-to
RAG Retrieval-Augmented Generation How-to
vLLM / SGLang / TensorRT-LLM Production serving engines How-to · versions
Benchmarks MMLU-Pro, GPQA, SWE-bench, LMArena, and more Reference
Structured Output Constrained decoding, JSON schema How-to
Safety / Guardrails Content filtering, prompt injection defense How-to · threat model
Glossary TTFT, prefill, decode, PagedAttention, and more Reference

Evaluation

Dimension Rating Notes
Maturity High The transformer has been battle-tested since 2017. MoE has dominated large open models since 2024–2025
Ecosystem Massive Hugging Face, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, MLX. Engines ship day-0 support for major open models
Accessibility Improving QLoRA fine-tunes 65–70B models on a single 48–80 GB GPU. Hosted APIs start at ~$1/MTok input
Local Inference Strong GGUF + llama.cpp/Ollama and MLX run 8B–32B models on consumer hardware, and large MoE models on 128–512 GB Macs
Rate of change Very high Model lineups, prices, and engine versions change monthly. Recheck Reference before relying on numbers

Pros

  • Open, well-documented architecture with mature open-source training and serving stacks
  • Strong open-weight options (DeepSeek, Qwen, gpt-oss, Llama, Gemma) under permissive licenses (MIT, Apache-2.0) or custom community licenses
  • Serving efficiency keeps compounding (PagedAttention, prefix caching, FP8/FP4, speculative decoding, disaggregation)

Cons

  • Memory (weights + KV cache) rather than compute is the usual bottleneck. MoE saves FLOPs, not VRAM
  • Quality after aggressive quantization and long-context behavior must be validated per model and task
  • Security is unsolved at the model level. Prompt injection needs defense-in-depth outside the model

Topic Map

  • How-to Guides: choose a serving engine (with a decision flowchart), estimate VRAM, pick GPUs and quantization formats, fine-tune with PEFT, scale out with Dynamo/llm-d, build RAG, enforce JSON schemas, configure guardrails, production checklist, and CLI recipes
  • Reference: model landscape and context windows, numeric formats, GGUF bits/weight, engine and tooling versions, GPU specs, benchmarks, vector databases, guardrail frameworks, OWASP LLM Top 10, and glossary
  • Explanation: transformer internals, attention variants, tokenization, training pipeline, MoE, quantization internals, inference optimization (serving architecture, KV cache, speculative decoding), MLX, distillation, scaling laws, merging, post-transformer architectures, post-training math, GPU parallelism, multimodality, and the LLM security threat model
  • Source notes: Ref: Alisa's Book of LLMs (derivation-heavy transformer, post-training, and parallelism notes) and Ref: Alisa's Math Notes (probability and statistics for ML)

Sources

Official Documentation and Repositories

Source Notes

Articles and Guides

Architecture Internals

Model Merging

VRAM & GPU

RAG

Benchmarks

Structured Output

Safety & Guardrails

Questions

  • How far will hybrid attention/SSM designs (Qwen3.5, Nemotron-H) replace full attention at frontier scale? See Explanation.
  • What is the practical floor for quantization (NVFP4, 2–3 bpw EXL3/IQ quants) before quality degrades unacceptably for agentic and tool-use workflows? See Explanation and Reference.
  • Disaggregated prefill/decode now ships in vLLM, SGLang, TensorRT-LLM, Dynamo, and llm-d. At what scale does it beat colocated serving with chunked prefill? See Explanation.
  • How will Apple's MLX ecosystem evolve with M5 Pro/Max memory and bandwidth (up to 128 GB and 614 GB/s on M5 Max, see Reference), and will an M5 Ultra follow?