Skip to content

AI Platform Engineering

Summary

AI Platform Engineering is the discipline of designing, building, and operating the infrastructure stack that takes machine learning models from experiments to reliable production services. Unlike traditional platform engineering (CPU-centric, database-backed), AI platform engineering centers on GPU compute, the most expensive and constrained resource in the stack. It extends Kubernetes with GPU drivers and device allocation (GPU Operator, device plugins, DRA), batch and gang schedulers (Kueue, Volcano, KAI Scheduler), distributed compute frameworks (Ray), and inference-optimized serving (vLLM, KServe, llm-d, NVIDIA Dynamo) to maximize GPU utilization while meeting latency SLAs.

Key Facts

Fact Value
Scope Kubernetes-based GPU infrastructure for training and LLM inference
Kubernetes baseline 1.37 (2026-08-26). DRA GA since 1.34 (2025-08), gang scheduling Workload API beta in 1.37
Latest Version (GPU node stack) NVIDIA GPU Operator v26.7.1 (2026-09-23)
Latest Version (schedulers) Kueue v0.19.6 (2026-09), KAI Scheduler v0.18.0 (2026-09-23), Volcano v1.15.2
Latest Version (compute and serving) Ray 2.58.0 (2026-08-23), KubeRay v1.7.1, vLLM 0.30.0 (2026-09-22), KServe 0.20.0 (2026-08-06), NVIDIA Dynamo 1.5.0 (2026-09-19), llm-d v0.7 (2026-05)
License All core projects Apache-2.0 (Run:ai is the main commercial layer)
Governance CNCF: Kubernetes (graduated), KServe, Volcano, Kubeflow (incubating), KAI Scheduler, llm-d (sandbox). Kueue and Gateway API Inference Extension are Kubernetes SIG subprojects
Primary cost driver GPU hours. Idle or fragmented GPUs are the main waste

Full version matrix with sources: Reference.

Why AI Workloads Are Different

Traditional applications follow a well-understood pattern: User → API → Database. The bottleneck is typically CPU, memory, or database performance, and horizontal scaling is straightforward.

AI inference workloads follow a fundamentally different pattern: User → Inference Server → GPU → Model Weights. The bottleneck shifts from data serving to compute serving:

Dimension Traditional Applications AI Applications
Primary resource CPU, memory GPU memory, GPU compute
Bottleneck Database I/O, network latency Model loading time, memory bandwidth
Idle cost Idle CPU is tolerable Idle GPU is extremely expensive
Scaling Horizontal (add replicas) GPU-aware (multi-GPU coordination)
Scheduling Place anywhere with capacity Topology-aware, gang scheduling
Serving Stateless request/response Stateful KV cache, batched inference, cache-aware routing

Core Problem Statement

The central challenge is GPU utilization. Organizations invest heavily in GPU hardware or cloud GPU hours to accelerate model execution. If those GPUs sit idle or are underutilized because of resource fragmentation, infrastructure costs increase without delivering value.

Platform engineers must answer:

  • How do we keep GPUs busy?
  • How do we schedule workloads without wasting compute?
  • How do we serve more inference requests per GPU?
  • How do we orchestrate distributed training across multiple GPUs and nodes?

Why Kubernetes Alone Is Not Enough

Kubernetes excels at scheduling containers against CPU and memory resources. AI workloads introduce requirements that standard Kubernetes has only recently started to address natively:

  • GPU-aware scheduling: GPUs must be discovered and registered via device plugins or DRA drivers before Kubernetes can manage them
  • Gang scheduling: distributed training jobs require all GPUs allocated simultaneously or not at all (native Workload API only reached beta in 1.37)
  • Topology awareness: GPU placement across nodes and racks affects inter-GPU communication latency
  • Multi-GPU coordination: training and inference can span multiple GPUs on a single node or across nodes
  • Resource sharing: time-slicing, MPS and MIG partitioning for efficient sub-GPU allocation
  • Distributed compute: task and actor scheduling above the infrastructure layer (Ray)
  • Inference optimization: KV cache management, continuous batching, PagedAttention (vLLM), and KV-cache-aware routing

This is why the AI infrastructure ecosystem developed tools like Kubeflow, KServe, Ray, vLLM, Volcano, Kueue, KAI Scheduler, llm-d and Dynamo. The reasoning behind each is in Explanation.

Platform at a Glance

The layers of a Kubernetes AI platform and the projects that occupy them (full diagram in Explanation):

flowchart TB
    subgraph SERVE["Serving"]
        GW["Gateway API + Inference Extension<br/>(EPP / llm-d router)"] --> SRV["KServe / llm-d / Dynamo / Ray Serve"]
        SRV --> ENG["vLLM / SGLang / TensorRT-LLM"]
    end
    subgraph TRAIN["Training and batch"]
        TJ["Kubeflow Trainer / RayJob / Volcano Job"] --> Q["Kueue or KAI queues"]
    end
    subgraph K8S["Kubernetes"]
        SCH["kube-scheduler / Volcano / KAI Scheduler"]
        ALLOC["Device plugin or DRA driver"]
    end
    subgraph NODE["GPU nodes"]
        OP["GPU Operator<br/>(driver, toolkit, DCGM, MIG Manager)"] --> GPU["NVIDIA GPUs<br/>(whole, MIG, time-sliced)"]
    end
    ENG --> SCH
    Q --> SCH
    SCH --> ALLOC --> GPU

The three-layer model (applications, MLOps services, infrastructure) is described in Explanation.

Tool Landscape

Layer Tool Purpose
Training Pipelines Kubeflow (Trainer v2, Pipelines), Argo Workflows End-to-end ML workflow orchestration
Model Registry MLflow Model versioning, tracking, and metadata
Distributed Compute Ray + KubeRay Task and actor scheduling across GPU clusters
Model Serving KServe, vLLM, llm-d, NVIDIA Dynamo, Ray Serve Production inference with GPU optimization
Inference Routing Gateway API Inference Extension, llm-d router KV-cache- and load-aware request routing
Batch Scheduling Kueue, Volcano, KAI Scheduler Gang scheduling, queue management, fair-share
GPU Management NVIDIA GPU Operator, DRA Driver for NVIDIA GPUs Driver lifecycle, device allocation, monitoring
Infrastructure Kubernetes Container orchestration and resource management

What Changed in 2025-2026

  • DRA went GA in Kubernetes 1.34 (2025-08). 1.36 (2026-04) made partitionable devices, consumable capacity and device taints beta. 1.37 (2026-08) made device taints and the extended-resource bridge GA, so DRA drivers can serve classic nvidia.com/gpu requests. (Reference)
  • Native gang scheduling: the Workload / PodGroup API is beta (v1beta1) in 1.37, with stable planned for 1.38.
  • NVIDIA donated its GPU DRA driver to the Kubernetes project at KubeCon Europe 2026. It now lives at kubernetes-sigs/dra-driver-nvidia-gpu. GPU allocation through it was still marked not officially supported as of 2026-09. ComputeDomains for GB200/GB300 are the mature part.
  • KAI Scheduler (open-sourced from Run:ai in 2025-04) became a CNCF sandbox project and adopted an LTS cadence (even minors, 1-year support).
  • Kueue moved to the v1beta2 API. v1beta1 support was dropped in v0.17.
  • LLM-aware routing standardized: the Gateway API Inference Extension reached GA with InferencePool v1. In 2026 its Endpoint Picker moved to llm-d/llm-d-router. llm-d joined the CNCF sandbox (2026-03). KServe added LLMInferenceService.
  • Disaggregated serving went mainstream: llm-d and NVIDIA Dynamo (1.x, Apache-2.0) both split prefill and decode and offload KV cache to CPU and SSD tiers.
  • Kubeflow Trainer v2 (TrainJob, 2025-07) replaced Training Operator v1, which is now in maintenance.
  • Ray's ray-ml images were discontinued. Use rayproject/ray:<ver>-gpu images.

Evaluation

Dimension Assessment
Maturity Consolidating: core APIs (DRA, InferencePool) are GA, but gang scheduling, DRA GPU drivers and disaggregated serving are still maturing
Complexity High: multi-layer stack with many moving parts
Cost Sensitivity Critical: GPU costs dominate infrastructure budgets
Ecosystem Very active: Kubernetes, Ray, vLLM, Kubeflow, KServe, llm-d, Dynamo ship monthly or faster
Entry Barrier Moderate to high: requires Kubernetes expertise plus GPU/ML domain knowledge
Pros Cons
Leverages existing Kubernetes skills Multi-layer complexity increases operational burden
Modular: components can be adopted incrementally GPU hardware costs remain significant
Active open-source ecosystem (CNCF, Ray, vLLM) Fast-moving ecosystem means frequent breaking changes (for example Kueue v1beta1 removal, ray-ml images)
Enables GPU sharing and cost optimization Requires specialized knowledge (GPU topology, distributed training)
Supports both training and inference workloads Several overlapping schedulers and serving stacks. Picking one is non-trivial

When it fits: you run GPUs you pay for continuously (on-prem or reserved cloud), have multiple teams competing for them, or serve self-hosted LLMs at more than a few replicas. When it does not: a single team with a few GPUs or bursty usage is usually better served by a managed inference API or a managed training service.

Topic Map

  • How-to Guides: prepare GPU nodes, allocate and share GPUs, schedule batch jobs, run Ray and vLLM, secure the platform, troubleshoot.
  • Reference: component versions, Kubernetes DRA and scheduling feature stages, MIG profiles, engine flags, metrics, ports, hardening checklists.
  • Explanation: device plugins vs DRA, GPU fragmentation and sharing, gang and topology-aware scheduling, the three-layer platform architecture.

Sources

Primary documentation and release sources (checked 2026-09-25):

Background reading used for the original structure of this topic:

Questions

  • What is the practical cost comparison between MIG partitioning, time-slicing, MPS and dedicated GPU allocation for inference workloads on current hardware (H200/B200)?
  • When does the NVIDIA DRA driver's GPU allocation become officially supported, and when should platforms move node pools from the device plugin to DRA using the 1.37 extended-resource bridge?
  • Kueue vs Volcano vs KAI Scheduler vs the native Workload API (beta in 1.37): which combination should a new platform standardize on for gang scheduling once KEP-4671 reaches GA?
  • llm-d vs NVIDIA Dynamo vs KServe LLMInferenceService: how much do they overlap after the Endpoint Picker moved to llm-d-router, and which becomes the default Kubernetes-native serving stack?
  • Will Ray remain the default distributed compute layer, or will Kubernetes-native primitives (JobSet, TrainJob, Workload API) absorb most of its infrastructure role?
  • How should platform teams approach GPU capacity planning when model sizes and inference patterns change rapidly?
  • What monitoring signals beyond standard latency/throughput are essential for GPU-intensive workloads (SM activity, Tensor Core activity, NVLink bandwidth, XID errors, KV cache hit rate)?