AI Platform Engineering¶
Summary
AI Platform Engineering is the discipline of designing, building, and operating the infrastructure stack that takes machine learning models from experiments to reliable production services. Unlike traditional platform engineering (CPU-centric, database-backed), AI platform engineering centers on GPU compute, the most expensive and constrained resource in the stack. It extends Kubernetes with GPU drivers and device allocation (GPU Operator, device plugins, DRA), batch and gang schedulers (Kueue, Volcano, KAI Scheduler), distributed compute frameworks (Ray), and inference-optimized serving (vLLM, KServe, llm-d, NVIDIA Dynamo) to maximize GPU utilization while meeting latency SLAs.
Key Facts¶
| Fact | Value |
|---|---|
| Scope | Kubernetes-based GPU infrastructure for training and LLM inference |
| Kubernetes baseline | 1.37 (2026-08-26). DRA GA since 1.34 (2025-08), gang scheduling Workload API beta in 1.37 |
| Latest Version (GPU node stack) | NVIDIA GPU Operator v26.7.1 (2026-09-23) |
| Latest Version (schedulers) | Kueue v0.19.6 (2026-09), KAI Scheduler v0.18.0 (2026-09-23), Volcano v1.15.2 |
| Latest Version (compute and serving) | Ray 2.58.0 (2026-08-23), KubeRay v1.7.1, vLLM 0.30.0 (2026-09-22), KServe 0.20.0 (2026-08-06), NVIDIA Dynamo 1.5.0 (2026-09-19), llm-d v0.7 (2026-05) |
| License | All core projects Apache-2.0 (Run:ai is the main commercial layer) |
| Governance | CNCF: Kubernetes (graduated), KServe, Volcano, Kubeflow (incubating), KAI Scheduler, llm-d (sandbox). Kueue and Gateway API Inference Extension are Kubernetes SIG subprojects |
| Primary cost driver | GPU hours. Idle or fragmented GPUs are the main waste |
Full version matrix with sources: Reference.
Why AI Workloads Are Different¶
Traditional applications follow a well-understood pattern: User → API → Database. The bottleneck is typically CPU, memory, or database performance, and horizontal scaling is straightforward.
AI inference workloads follow a fundamentally different pattern: User → Inference Server → GPU → Model Weights. The bottleneck shifts from data serving to compute serving:
| Dimension | Traditional Applications | AI Applications |
|---|---|---|
| Primary resource | CPU, memory | GPU memory, GPU compute |
| Bottleneck | Database I/O, network latency | Model loading time, memory bandwidth |
| Idle cost | Idle CPU is tolerable | Idle GPU is extremely expensive |
| Scaling | Horizontal (add replicas) | GPU-aware (multi-GPU coordination) |
| Scheduling | Place anywhere with capacity | Topology-aware, gang scheduling |
| Serving | Stateless request/response | Stateful KV cache, batched inference, cache-aware routing |
Core Problem Statement¶
The central challenge is GPU utilization. Organizations invest heavily in GPU hardware or cloud GPU hours to accelerate model execution. If those GPUs sit idle or are underutilized because of resource fragmentation, infrastructure costs increase without delivering value.
Platform engineers must answer:
- How do we keep GPUs busy?
- How do we schedule workloads without wasting compute?
- How do we serve more inference requests per GPU?
- How do we orchestrate distributed training across multiple GPUs and nodes?
Why Kubernetes Alone Is Not Enough¶
Kubernetes excels at scheduling containers against CPU and memory resources. AI workloads introduce requirements that standard Kubernetes has only recently started to address natively:
- GPU-aware scheduling: GPUs must be discovered and registered via device plugins or DRA drivers before Kubernetes can manage them
- Gang scheduling: distributed training jobs require all GPUs allocated simultaneously or not at all (native Workload API only reached beta in 1.37)
- Topology awareness: GPU placement across nodes and racks affects inter-GPU communication latency
- Multi-GPU coordination: training and inference can span multiple GPUs on a single node or across nodes
- Resource sharing: time-slicing, MPS and MIG partitioning for efficient sub-GPU allocation
- Distributed compute: task and actor scheduling above the infrastructure layer (Ray)
- Inference optimization: KV cache management, continuous batching, PagedAttention (vLLM), and KV-cache-aware routing
This is why the AI infrastructure ecosystem developed tools like Kubeflow, KServe, Ray, vLLM, Volcano, Kueue, KAI Scheduler, llm-d and Dynamo. The reasoning behind each is in Explanation.
Platform at a Glance¶
The layers of a Kubernetes AI platform and the projects that occupy them (full diagram in Explanation):
flowchart TB
subgraph SERVE["Serving"]
GW["Gateway API + Inference Extension<br/>(EPP / llm-d router)"] --> SRV["KServe / llm-d / Dynamo / Ray Serve"]
SRV --> ENG["vLLM / SGLang / TensorRT-LLM"]
end
subgraph TRAIN["Training and batch"]
TJ["Kubeflow Trainer / RayJob / Volcano Job"] --> Q["Kueue or KAI queues"]
end
subgraph K8S["Kubernetes"]
SCH["kube-scheduler / Volcano / KAI Scheduler"]
ALLOC["Device plugin or DRA driver"]
end
subgraph NODE["GPU nodes"]
OP["GPU Operator<br/>(driver, toolkit, DCGM, MIG Manager)"] --> GPU["NVIDIA GPUs<br/>(whole, MIG, time-sliced)"]
end
ENG --> SCH
Q --> SCH
SCH --> ALLOC --> GPU
The three-layer model (applications, MLOps services, infrastructure) is described in Explanation.
Tool Landscape¶
| Layer | Tool | Purpose |
|---|---|---|
| Training Pipelines | Kubeflow (Trainer v2, Pipelines), Argo Workflows | End-to-end ML workflow orchestration |
| Model Registry | MLflow | Model versioning, tracking, and metadata |
| Distributed Compute | Ray + KubeRay | Task and actor scheduling across GPU clusters |
| Model Serving | KServe, vLLM, llm-d, NVIDIA Dynamo, Ray Serve | Production inference with GPU optimization |
| Inference Routing | Gateway API Inference Extension, llm-d router | KV-cache- and load-aware request routing |
| Batch Scheduling | Kueue, Volcano, KAI Scheduler | Gang scheduling, queue management, fair-share |
| GPU Management | NVIDIA GPU Operator, DRA Driver for NVIDIA GPUs | Driver lifecycle, device allocation, monitoring |
| Infrastructure | Kubernetes | Container orchestration and resource management |
What Changed in 2025-2026¶
- DRA went GA in Kubernetes 1.34 (2025-08). 1.36 (2026-04) made partitionable devices, consumable capacity and device taints beta. 1.37 (2026-08) made device taints and the extended-resource bridge GA, so DRA drivers can serve classic
nvidia.com/gpurequests. (Reference) - Native gang scheduling: the Workload / PodGroup API is beta (
v1beta1) in 1.37, with stable planned for 1.38. - NVIDIA donated its GPU DRA driver to the Kubernetes project at KubeCon Europe 2026. It now lives at
kubernetes-sigs/dra-driver-nvidia-gpu. GPU allocation through it was still marked not officially supported as of 2026-09. ComputeDomains for GB200/GB300 are the mature part. - KAI Scheduler (open-sourced from Run:ai in 2025-04) became a CNCF sandbox project and adopted an LTS cadence (even minors, 1-year support).
- Kueue moved to the
v1beta2API.v1beta1support was dropped in v0.17. - LLM-aware routing standardized: the Gateway API Inference Extension reached GA with
InferencePoolv1. In 2026 its Endpoint Picker moved tollm-d/llm-d-router. llm-d joined the CNCF sandbox (2026-03). KServe addedLLMInferenceService. - Disaggregated serving went mainstream: llm-d and NVIDIA Dynamo (1.x, Apache-2.0) both split prefill and decode and offload KV cache to CPU and SSD tiers.
- Kubeflow Trainer v2 (TrainJob, 2025-07) replaced Training Operator v1, which is now in maintenance.
- Ray's
ray-mlimages were discontinued. Userayproject/ray:<ver>-gpuimages.
Evaluation¶
| Dimension | Assessment |
|---|---|
| Maturity | Consolidating: core APIs (DRA, InferencePool) are GA, but gang scheduling, DRA GPU drivers and disaggregated serving are still maturing |
| Complexity | High: multi-layer stack with many moving parts |
| Cost Sensitivity | Critical: GPU costs dominate infrastructure budgets |
| Ecosystem | Very active: Kubernetes, Ray, vLLM, Kubeflow, KServe, llm-d, Dynamo ship monthly or faster |
| Entry Barrier | Moderate to high: requires Kubernetes expertise plus GPU/ML domain knowledge |
| Pros | Cons |
|---|---|
| Leverages existing Kubernetes skills | Multi-layer complexity increases operational burden |
| Modular: components can be adopted incrementally | GPU hardware costs remain significant |
| Active open-source ecosystem (CNCF, Ray, vLLM) | Fast-moving ecosystem means frequent breaking changes (for example Kueue v1beta1 removal, ray-ml images) |
| Enables GPU sharing and cost optimization | Requires specialized knowledge (GPU topology, distributed training) |
| Supports both training and inference workloads | Several overlapping schedulers and serving stacks. Picking one is non-trivial |
When it fits: you run GPUs you pay for continuously (on-prem or reserved cloud), have multiple teams competing for them, or serve self-hosted LLMs at more than a few replicas. When it does not: a single team with a few GPUs or bursty usage is usually better served by a managed inference API or a managed training service.
Topic Map¶
- How-to Guides: prepare GPU nodes, allocate and share GPUs, schedule batch jobs, run Ray and vLLM, secure the platform, troubleshoot.
- Reference: component versions, Kubernetes DRA and scheduling feature stages, MIG profiles, engine flags, metrics, ports, hardening checklists.
- Explanation: device plugins vs DRA, GPU fragmentation and sharing, gang and topology-aware scheduling, the three-layer platform architecture.
Related Topics¶
- Kubernetes: container orchestration fundamentals
- Docker: container runtime
- LLM Fundamentals: model architecture, attention, KV cache, quantization
- LLM Inference: speculative decoding and engine-level serving optimizations (SGLang, vLLM, TensorRT-LLM)
- DFlash 2: speculative-decoding drafter that runs inside vLLM and SGLang
- OpenTelemetry: tracing and metrics conventions for inference services
- Infrastructure comparisons: domain-level platform comparisons
Sources¶
Primary documentation and release sources (checked 2026-09-25):
- Kubernetes v1.37 release, v1.37 DRA updates, v1.36 DRA updates, v1.34 DRA GA
- Kubernetes Dynamic Resource Allocation and Device Plugins
- NVIDIA GPU Operator release notes and v26.7.1 release
- NVIDIA k8s-device-plugin, DCGM Exporter
- DRA Driver for NVIDIA GPUs and NVIDIA at KubeCon 2026
- NVIDIA MIG User Guide, Run:ai GPU Time-Slicing
- Kueue, Volcano (CNCF project page), KAI Scheduler (CNCF project page)
- Kubeflow Trainer
- Ray on PyPI, Ray Cluster Key Concepts, Ray on Kubernetes, KubeRay, KubeRay v1.5 announcement
- vLLM (PyPI), vLLM production-stack, PagedAttention paper (arXiv:2309.06180)
- KServe, Gateway API Inference Extension, llm-d (CNCF welcome post), NVIDIA Dynamo
Background reading used for the original structure of this topic:
- 7 Days of AI Platform Engineering — Milind Dethe
- Day 1: Why AI Workloads Are Different
- Day 2: GPU Discovery in Kubernetes
- Day 3: GPU Utilization & Resource Fragmentation
- Day 4: AI Workload Scheduling
- Day 5: Ray — Distributed Compute
- Day 6: vLLM — Efficient LLM Serving
- Day 7: The Full AI Platform Stack
- vLLM website, KV Caching Explained — Hugging Face
Questions¶
- What is the practical cost comparison between MIG partitioning, time-slicing, MPS and dedicated GPU allocation for inference workloads on current hardware (H200/B200)?
- When does the NVIDIA DRA driver's GPU allocation become officially supported, and when should platforms move node pools from the device plugin to DRA using the 1.37 extended-resource bridge?
- Kueue vs Volcano vs KAI Scheduler vs the native Workload API (beta in 1.37): which combination should a new platform standardize on for gang scheduling once KEP-4671 reaches GA?
- llm-d vs NVIDIA Dynamo vs KServe
LLMInferenceService: how much do they overlap after the Endpoint Picker moved tollm-d-router, and which becomes the default Kubernetes-native serving stack? - Will Ray remain the default distributed compute layer, or will Kubernetes-native primitives (JobSet, TrainJob, Workload API) absorb most of its infrastructure role?
- How should platform teams approach GPU capacity planning when model sizes and inference patterns change rapidly?
- What monitoring signals beyond standard latency/throughput are essential for GPU-intensive workloads (SM activity, Tensor Core activity, NVLink bandwidth, XID errors, KV cache hit rate)?