Skip to content

AI Platform Engineering — Explanation

Why the AI platform stack looks the way it does: how Kubernetes learns about GPUs (device plugins, then DRA), why GPUs fragment and how sharing works, why AI jobs need gang and topology-aware scheduling, how Ray and vLLM schedule work above Kubernetes, and how the LLM-aware serving layer (KServe, Gateway API Inference Extension, llm-d, NVIDIA Dynamo) routes requests by KV cache. Versions and look-up tables live in Reference. Recipes live in How-to Guides.

The Three-Layer AI Platform Architecture

A production AI platform consists of three distinct layers:

Layer 1: AI Applications (Business Layer)

The products users interact with:

  • Virtual assistants, chatbots and agents
  • Recommendation systems
  • Fraud detection pipelines
  • Content generation services
  • IoT analytics

Layer 2: AI Platform Services (MLOps Layer)

The platform that takes raw data and turns it into production models:

Stage Purpose Example Tools
Data Processing Collect, clean, label, store data Spark, Flink, Ray Data, Label Studio
Feature Engineering Transform raw data into model features Feature Stores (Feast, Tecton)
Model Training Distributed GPU training Kubeflow Trainer (TrainJob), Ray Train
Model Registry Version, store, track models MLflow, Weights & Biases
Deployment & Inference Serve predictions reliably KServe, Ray Serve, vLLM, llm-d, NVIDIA Dynamo
Monitoring Track latency, drift, costs Prometheus, Grafana, DCGM Exporter, Evidently

Layer 3: Infrastructure (Compute Layer)

The foundation everything depends on:

Category Components
Compute CPUs, GPUs, TPUs
Storage Object storage, data lakes, feature stores
Orchestration Kubernetes, networking, scheduling (Kueue, Volcano, KAI Scheduler)
Accelerators NVIDIA GPUs (A100, H100/H200, B200/GB200), AMD Instinct, Google TPUs

GPU Discovery in Kubernetes

Kubernetes does not natively understand GPU hardware. The classic discovery path makes GPUs visible as schedulable extended resources:

flowchart TB
    HW["GPU hardware<br/>(PCIe / SXM)"] --> DRV["NVIDIA driver<br/>(kernel module)"]
    DRV --> CTK["NVIDIA Container Toolkit<br/>(CDI spec)"]
    CTK --> DP["NVIDIA device plugin<br/>(DaemonSet)"]
    DP -->|"gRPC ListAndWatch"| KL["kubelet<br/>node.status.allocatable"]
    KL --> SCH["kube-scheduler<br/>matches nvidia.com/gpu requests"]
    SCH --> POD["Pod with GPU devices injected"]

Step 1: Physical GPU + Driver

The physical GPU is attached to a worker node. The operating system exposes it through vendor-specific drivers. For NVIDIA GPUs, the NVIDIA driver must be installed on the node before the OS can interact with the hardware.

Step 2: Kubernetes Awareness Gap

Even with the driver installed, Kubernetes remains unaware of the GPU. The built-in resources the scheduler counts are:

  • CPU (cpu)
  • Memory (memory)
  • Ephemeral storage (ephemeral-storage)
  • Huge pages (hugepages-<size>)

GPUs must be explicitly registered, either through the Device Plugin API or through DRA.

Step 3: Device Plugin Registration

A Device Plugin is a Kubernetes extension that advertises specialized hardware to the kubelet. The NVIDIA Device Plugin runs as a DaemonSet on each GPU node and reports available GPUs.

Once registered, Kubernetes sees the resource:

# Node capacity after device plugin registration
nvidia.com/gpu: 4

Key Insight

Kubernetes does not schedule GPUs because it understands GPUs. It schedules GPUs because a Device Plugin exposes them as generic extended resources, which are opaque integer counters. The same mechanism works for TPUs, FPGAs, SmartNICs, and any other accelerator with a Device Plugin implementation. That opacity is the limitation DRA was built to remove.

Step 4: Pod GPU Requests

After registration, workloads request GPUs identically to CPU and memory:

apiVersion: v1
kind: Pod
metadata:
  name: gpu-inference
spec:
  containers:
    - name: model-server
      image: vllm/vllm-openai:v0.30.0
      resources:
        limits:
          nvidia.com/gpu: 1

The scheduler matches resource requests with available resources. It has no knowledge of whether the workload is running an LLM, image generator, or training job, nor of GPU model, memory size or NVLink topology beyond what node labels say.

Device Plugin Architecture

The registration and allocation handshake between the device plugin, kubelet and scheduler:

sequenceDiagram
    participant GPU as GPU Hardware
    participant Driver as NVIDIA Driver
    participant DP as NVIDIA Device Plugin
    participant Kubelet as kubelet
    participant Scheduler as kube-scheduler
    participant Pod as Pod

    GPU->>Driver: Hardware attached
    Driver->>DP: GPU devices available (NVML)
    DP->>Kubelet: Register via gRPC<br/>ListAndWatch(nvidia.com/gpu: 4)
    Kubelet->>Scheduler: Node allocatable updated
    Pod->>Scheduler: Request nvidia.com/gpu: 1
    Scheduler->>Kubelet: Bind Pod to GPU node
    Kubelet->>DP: Allocate(deviceID)
    DP->>Pod: CDI device names / device nodes + env vars

The Device Plugin communicates with the kubelet via a gRPC interface over Unix sockets in /var/lib/kubelet/device-plugins/. The plugin must handle kubelet restarts by monitoring socket deletion and re-registering. The Device Plugin API supports:

  • ListAndWatch — advertises available devices and reports health changes
  • Allocate — provisions device access for containers (device nodes, environment variables, mounts, CDI device names)
  • Health monitoring — marks devices as unhealthy when failures are detected. This reduces the node's allocatable count

Since Kubernetes v1.36 (beta, KEP-4680), allocatedResourcesStatus in pod status reports per-device health for both device plugins and DRA drivers.

NVIDIA GPU Operator

The NVIDIA GPU Operator packages the whole node stack (containerized driver, Container Toolkit, device plugin, GPU Feature Discovery, DCGM Exporter, MIG Manager) as operands of one ClusterPolicy custom resource. It removes the need for a special GPU OS image: a new node joins, NFD labels it as NVIDIA hardware, and the operator rolls the stack onto it. The component list and versions per release are in Reference.


Dynamic Resource Allocation (DRA)

DRA (resource.k8s.io/v1, GA since Kubernetes 1.34) replaces "count of opaque devices" with structured device descriptions the scheduler can reason about. A DRA driver publishes each node's devices and their attributes (model, memory, MIG capability, NVLink/PCIe topology) as ResourceSlice objects. Workloads create ResourceClaims, directly or via ResourceClaimTemplates, that select devices from a DeviceClass with CEL expressions. The scheduler allocates concrete devices before binding the pod. The kubelet plugin then prepares them (via CDI) at pod start.

How a GPU claim flows through DRA, from driver inventory to container start:

sequenceDiagram
    participant Drv as DRA kubelet plugin<br/>(gpu.nvidia.com)
    participant API as kube-apiserver
    participant User as Workload owner
    participant Sch as kube-scheduler<br/>(DRA plugin)
    participant Kl as kubelet
    Drv->>API: Publish ResourceSlice<br/>(GPUs + attributes per node)
    User->>API: Create Pod + ResourceClaimTemplate<br/>(deviceClassName gpu.nvidia.com)
    API->>API: Generate ResourceClaim for the Pod
    Sch->>API: Read ResourceSlices, evaluate CEL selectors
    Sch->>API: Write allocation to ResourceClaim status, bind Pod
    Kl->>Drv: NodePrepareResources(claim)
    Drv->>Kl: CDI device IDs (GPU, MIG slice, or IMEX channel)
    Kl->>Kl: Start containers with devices injected

Why this matters for AI platforms:

Device plugin limitation DRA answer
Only integer counts (nvidia.com/gpu: 1) Requests by attribute: "one GPU with at least 80 GiB", "an H100 or else two A100s" (prioritized list, GA in 1.36)
MIG layouts must be pre-sliced on the node Partitionable devices (beta in 1.36) let the scheduler pick a partition size at scheduling time
Sharing needs per-plugin hacks One ResourceClaim can be referenced by several containers or pods
No way to drain a bad GPU gracefully Device taints and tolerations (GA in 1.37)
Migration requires rewriting every manifest Extended resource mapping (GA in 1.37) lets a DRA driver satisfy classic nvidia.com/gpu requests
Multi-node NVLink is invisible NVIDIA ComputeDomains allocate IMEX domains across GB200/GB300 NVL72 nodes

NVIDIA donated its DRA driver to the Kubernetes project at KubeCon Europe 2026. It now lives at kubernetes-sigs/dra-driver-nvidia-gpu and ships from registry.k8s.io. Its ComputeDomain side is production-ready. GPU allocation was still flagged "not yet officially supported" in the README as of 2026-09, which is why most clusters still run the classic device plugin for plain GPUs.

Device plugins are not deprecated

DRA and device plugins coexist. The extended-resource bridge exists precisely so clusters can move node pools to DRA without changing tenant manifests. Expect a multi-year transition.


GPU Utilization and Resource Fragmentation

The Core Problem

Standard Kubernetes GPU allocation is binary. A pod requests one or more whole GPUs, and each GPU is allocated exclusively:

Pod A → 1 GPU Requested → 1 GPU Allocated (80 GB)
Actual usage: 10 GB memory, 20% compute
Waste: 70 GB memory, 80% compute

With CPUs, Kubernetes efficiently bin-packs multiple pods onto a single node. With the device plugin model, GPUs do not support this: one pod per GPU, regardless of actual utilization.

Cost Impact

Consider a cluster with 8x NVIDIA A100 GPUs where every workload uses only 25% of each GPU:

  • Effective utilization: 2 GPUs worth of useful work
  • Paid for: 8 GPUs
  • Waste: 75% of GPU investment

At an illustrative purchase price of about $30,000 per A100 (2023-era list pricing; street and cloud prices vary widely and have fallen since), the six idle GPUs represent about $180,000 of stranded capacity.

GPU Sharing Strategies

Four techniques address GPU underutilization at the infrastructure level, plus one at the serving layer. A side-by-side table is in Reference.

Time-Slicing

Multiple workloads take turns using the same GPU, analogous to CPU time-sharing. The NVIDIA device plugin's native time-slicing simply advertises each GPU N times (replicas). The CUDA driver round-robins contexts with no memory limits and no fault isolation.

NVIDIA Run:ai adds scheduler-controlled time-slicing with two modes:

Mode Behavior K8s Mapping
Strict Each workload gets exactly its requested GPU compute fraction gpu-compute-request = gpu-compute-limit = gpu-fraction
Fair Each workload gets at least its fraction, plus unused slices from idle workloads gpu-compute-request = gpu-fraction, gpu-compute-limit = 1.0

Per the Run:ai documentation, time-slicing operates on a plan/lease cycle. Default configuration:

  • Lease time: 250ms (exclusive GPU access per workload)
  • Granularity: 5% precision
  • Plan (cycle) time: 250ms / 0.05 = 5000ms (5 seconds)

A workload requesting gpu-fraction=0.5 gets 2.5s of runtime per 5s cycle.

Trade-offs

Decreasing lease time makes time-slicing less accurate. Increasing lease time improves accuracy but reduces workload responsiveness. Context switching between workloads adds overhead.

Multi-Process Service (MPS)

CUDA MPS runs kernels from several processes concurrently on one GPU through a shared control daemon, instead of alternating contexts. The NVIDIA device plugin supports sharing.mps, which also sets per-client memory and thread limits. MPS gives better throughput than time-slicing for many small inference processes, but clients still share one fault domain.

Multi-Instance GPU (MIG)

Available on data-center GPUs from Ampere onward (A30, A100, H100/H200, GH200, B200/GB200, RTX PRO 6000 Blackwell). MIG partitions a single GPU into up to 7 isolated GPU Instances, each with dedicated:

  • Streaming Multiprocessors (SMs)
  • GPU engines (copy engines, decoders)
  • L2 cache banks
  • Memory controllers
  • DRAM address busses

Example partitioning of an 80 GB A100:

Full A100 (80 GB)
├── MIG Instance 1: 10 GB (1g.10gb)
├── MIG Instance 2: 10 GB (1g.10gb)
├── MIG Instance 3: 20 GB (2g.20gb)
└── MIG Instance 4: 40 GB (4g.40gb)

The example is illustrative. Valid combinations are constrained by slice placement, so check with nvidia-smi mig -lgipp. Each instance provides hardware-level isolation: one workload cannot impact the L2 cache or DRAM bandwidth of another. This makes MIG suitable for multi-tenant environments where QoS guarantees are required. The cost is rigidity: changing a layout requires draining the GPU, which DRA partitionable devices aim to automate.

MIG supports:

  • Bare-metal and containers
  • GPU passthrough virtualization
  • vGPU on supported hypervisors

Continuous Batching (Inference)

Instead of processing inference requests one-by-one, the serving engine combines multiple requests into a single GPU execution. vLLM's continuous batching dynamically adds new requests as older ones complete. The GPU then stays busy continuously instead of waiting for a full batch to form. For LLM serving this usually recovers more utilization than any infrastructure-level sharing.


AI Workload Scheduling

Why Standard Kubernetes Scheduling Falls Short

Traditional applications are loosely coupled: components can start independently and tolerate staggered scheduling. The default kube-scheduler places one pod at a time. AI training jobs have fundamentally different requirements.

Gang Scheduling

Distributed training jobs require all resources allocated simultaneously:

Training Job (requires 8 GPUs)
├── Worker 0: GPU 0
├── Worker 1: GPU 1
├── Worker 2: GPU 2
├── Worker 3: GPU 3
├── Worker 4: GPU 4
├── Worker 5: GPU 5
├── Worker 6: GPU 6
└── Worker 7: GPU 7

If only 6 GPUs are available, the job cannot start. Partial allocation wastes resources: workers wait indefinitely for the remaining GPUs, blocking other jobs, and two half-placed jobs can deadlock each other.

Gang scheduling rule: Either schedule all required resources together, or schedule none of them.

Until recently this required a replacement scheduler (Volcano, KAI) or Kueue's admission-level approximation. Kubernetes added a native Workload / PodGroup API (KEP-4671): alpha in 1.35, beta (v1beta1) in 1.37 with a hierarchical CompositePodGroup alpha, and stable planned for 1.38.

Topology Awareness

GPU placement affects training performance significantly. Same-node GPUs communicate via NVLink (900 GB/s per GPU on H100), while cross-node GPUs use network fabric (InfiniBand NDR at 400 Gb/s, about 50 GB/s, per port). Bandwidth figures are in Reference.

The two tiers of GPU connectivity a topology-aware scheduler has to respect:

graph LR
    subgraph "Node A (Fast: NVLink)"
        GPU0["GPU 0"]
        GPU1["GPU 1"]
        GPU2["GPU 2"]
        GPU3["GPU 3"]
        GPU0 <--> GPU1
        GPU1 <--> GPU2
        GPU2 <--> GPU3
    end

    subgraph "Node B (Fast: NVLink)"
        GPU4["GPU 4"]
        GPU5["GPU 5"]
        GPU6["GPU 6"]
        GPU7["GPU 7"]
        GPU4 <--> GPU5
        GPU5 <--> GPU6
        GPU6 <--> GPU7
    end

    GPU3 <-.->|"Slower: InfiniBand / RoCE"| GPU4

A topology-aware scheduler prefers placing all GPUs on the same node when possible. If that is not possible, it falls back to nodes under the same leaf switch or rack (Kueue TAS and KAI TAS read Topology levels from node labels; Volcano uses HyperNodes). Rack-scale systems such as GB200 NVL72 add a third tier: an NVLink domain spanning 72 GPUs across nodes, exposed through DRA ComputeDomains.

Scheduling Tools Comparison

Tool Type Key Capabilities
Kueue Kubernetes SIG job queueing Admission control, quotas with borrowing, fair sharing, TAS, MultiKueue. Keeps kube-scheduler
Volcano CNCF incubating batch scheduler Gang scheduling, queues, DRF, network-topology-aware placement. v1.14 added a sharding controller for mixed agent/batch clusters
KAI Scheduler CNCF sandbox AI scheduler (from Run:ai) Hierarchical queues, gang and hierarchical PodGroups, GPU fractions, time-based fairshare, DRA, TAS
Kubeflow Trainer Training operator v2 TrainJob + runtimes on JobSet. Integrates with Kueue, Volcano and KAI for gang/TAS. Training Operator v1 (PyTorchJob etc.) is in maintenance
NVIDIA GPU Operator GPU lifecycle manager Driver management, device plugins, DCGM metrics, MIG management

The pattern most platforms converge on: Kueue (or KAI's queues) decides whether and when a job may run against a quota. A gang-capable scheduler decides where its pods land. A feature-by-feature matrix is in Reference.


Ray — Distributed Compute Framework

The Two-Layer Model

Ray schedules computation inside pods that Kubernetes has already placed:

graph TB
    subgraph "Application Layer"
        User["Python program"]
        RayDriver["Ray Driver<br/>(ray.init)"]
    end

    subgraph "Ray Layer (Computation Scheduling)"
        RayHead["Ray Head Node<br/>(GCS, Autoscaler, Dashboard)"]
        RayWorker1["Ray Worker 1<br/>(raylet + object store)"]
        RayWorker2["Ray Worker 2<br/>(raylet + object store)"]
        RayWorker3["Ray Worker 3<br/>(raylet + object store)"]
    end

    subgraph "Kubernetes Layer (Infrastructure Scheduling)"
        KubeRay["KubeRay operator<br/>(RayCluster / RayJob / RayService)"]
        K8s["kube-scheduler"]
        Pod1["Pod (Head)"]
        Pod2["Pod (Worker)"]
        Pod3["Pod (Worker)"]
        Pod4["Pod (Worker)"]
    end

    User --> RayDriver
    RayDriver --> RayHead
    RayHead --> RayWorker1
    RayHead --> RayWorker2
    RayHead --> RayWorker3

    KubeRay --> K8s
    K8s --> Pod1
    K8s --> Pod2
    K8s --> Pod3
    K8s --> Pod4

    Pod1 -.- RayHead
    Pod2 -.- RayWorker1
    Pod3 -.- RayWorker2
    Pod4 -.- RayWorker3

Kubernetes schedules infrastructure (which node should this pod run on?). Ray schedules computation (which worker executes which task? how are results collected?).

Ray Architecture

Component Role
Head Node Runs GCS (Global Control Service), autoscaler, Ray dashboard. Also schedules tasks like worker nodes unless its num-gpus/num-cpus are set to 0.
Worker Nodes Execute Ray tasks and actors. Each runs a raylet and participates in the distributed object store.
Autoscaler Scales worker nodes based on task/actor resource requests (not CPU/memory metrics).
GCS Central metadata store for cluster state, actor locations, and resource availability.

Tasks vs Actors

Dimension Tasks Actors
State Stateless Stateful
Lifecycle Run once, return result Long-lived, handle multiple requests
Use case Data processing, hyperparameter search Model serving, stateful computation
Invocation function.remote(args) actor.method.remote(args)

Tasks enable embarrassingly parallel workloads (data processing, hyperparameter tuning). Actors enable stateful services (model serving, game environments, RL training).

Ray Libraries

Library Purpose
Ray Train Distributed training (PyTorch, TensorFlow, XGBoost)
Ray Tune Hyperparameter optimization
Ray Serve Model serving and composition, including Ray Serve LLM on vLLM
Ray Data Distributed data processing, including batch LLM inference
Ray RLlib Reinforcement learning

Ray on Kubernetes (KubeRay)

KubeRay provides Kubernetes CRDs for managing Ray clusters:

  • RayCluster — manages head and worker pods
  • RayJob — creates a cluster (or uses an existing one), submits an entrypoint, and can tear the cluster down afterwards
  • RayService — manages Ray Serve deployments with zero-downtime upgrades (incremental upgrades since v1.5)

With enableInTreeAutoscaling: true, KubeRay injects the Ray autoscaler as a sidecar container in the head pod. It scales worker pods between minReplicas and maxReplicas based on pending Ray task/actor resource demands. KubeRay v1.5 added Ray token authentication support. KubeRay integrates with Kueue, Volcano and KAI Scheduler for gang admission of whole Ray clusters.


vLLM — Inference Engine Architecture

Why Naive Model Serving Fails at Scale

A naive inference server processes requests sequentially or in static batches:

Request 1 → GPU → Response 1
Request 2 → GPU → Response 2  (waits for Request 1)
Request 3 → GPU → Response 3  (waits for Request 2)

With 100 concurrent users, GPU utilization stays low because the GPU cannot exploit its parallel architecture. LLM decoding is memory-bandwidth-bound: every generated token re-reads all model weights, so throughput comes from batching many sequences into each step. This is the GPU utilization problem applied to inference serving.

PagedAttention

The key innovation in vLLM. Traditional serving systems allocate GPU memory for the KV cache in large, contiguous chunks. For variable-length sequences, this causes:

  • Internal fragmentation — allocated blocks larger than needed
  • External fragmentation — free memory scattered in unusable small chunks
  • Reservation waste — memory reserved for maximum sequence length even for short sequences

PagedAttention treats the KV cache like virtual memory. Memory is managed in fixed-size blocks (pages) that need not be contiguous, and a block table maps each sequence's logical blocks to physical ones:

Traditional KV Cache:
[████████████░░░░░░░░] Request 1 (wasted space)
[████████░░░░░░░░░░░░] Request 2 (wasted space)
[░░░░░░░░░░░░░░░░░░░░] Free (fragmented)

PagedAttention:
[████][████][████][██] Request 1 (pages, no waste)
[████][████][██]       Request 2 (pages, no waste)
[████][████]           Free (reusable pages)

Results reported in the PagedAttention paper (arXiv:2309.06180):

  • Existing systems wasted 60-80% of KV cache memory. vLLM keeps waste under 4% (only the last partial block)
  • 2-4x higher throughput than FasterTransformer and Orca at the same latency
  • Block-level sharing enables copy-on-write for parallel sampling and beam search, and later automatic prefix caching

Continuous Batching

Traditional batching waits for a full batch before processing:

Traditional:    Wait → Process Batch → Wait → Process Batch
Continuous:     Process ─── Process ─── Process ─── Process
                (new requests added as old ones complete)

vLLM's continuous (iteration-level) batching adds new requests to the running batch as existing requests finish generating tokens. The GPU stays busy continuously. The V1 engine (the only engine in current releases) also uses chunked prefill: long prompts are split so prefill chunks and decode tokens share each step under a --max-num-batched-tokens budget.

vLLM in the Platform Stack

Users
  |
vLLM (inference optimization: PagedAttention + continuous batching)
  |
Model Weights (loaded into GPU memory)
  |
GPU (compute)

In a Kubernetes deployment:

Users
  |
Gateway (Gateway API + Inference Extension / AI gateway)
  |
vLLM Pods (plain Deployment, KServe LLMInferenceService, llm-d, Dynamo, or Ray Serve)
  |
Kubernetes (scheduling, scaling, health checks)
  |
GPU Nodes (NVIDIA Device Plugin or DRA driver, GPU Operator)

vLLM exposes an OpenAI-compatible API. It is a drop-in replacement for existing applications that use the OpenAI API format.


LLM-Aware Serving Layer

Once an LLM runs on more than a handful of replicas, the load balancer becomes a performance component. Round-robin ignores which replica already holds a request's prefix in KV cache, how deep each replica's queue is, and which LoRA adapters are loaded. Since 2025 a Kubernetes-native layer has formed to fix that.

Project What it adds Relationship
Gateway API Inference Extension (GAIE) InferencePool API (GA, inference.networking.k8s.io/v1) and the ext-proc Endpoint Picker (EPP) protocol that turns Envoy Gateway, kgateway, Istio or GKE Gateway into an "inference gateway" Standard API. As of 2026 the full EPP, InferenceObjective and body-based router moved to llm-d/llm-d-router. GAIE keeps InferencePool, a lightweight reference EPP and conformance tests
llm-d Well-lit paths on top of vLLM: prefix-cache and load-aware routing, tiered KV cache offload, prefill/decode disaggregation, wide expert parallelism CNCF sandbox (2026-03), founded by Red Hat, Google Cloud, IBM Research, CoreWeave and NVIDIA
KServe LLMInferenceService CRD (router, prefill, worker, parallelism, KV offload) alongside the classic InferenceService CNCF incubating. Uses llm-d scheduling and GAIE resources under the hood
NVIDIA Dynamo Engine-agnostic (vLLM, SGLang, TensorRT-LLM) frontend, KV-aware router, SLA-based Planner autoscaler, KV Block Manager (GPU → CPU → SSD → remote), NIXL transfers, Grove gang-scheduling operator Apache-2.0. Can run its own frontend or plug into GAIE as an EPP
vLLM production-stack Helm chart with a session/prefix-aware router and LMCache offload Lighter-weight reference stack from the vLLM project

The request path through an inference gateway with an Endpoint Picker, which is the common shape across GAIE, llm-d and KServe:

sequenceDiagram
    participant C as Client (OpenAI API)
    participant GW as Gateway<br/>(Envoy Gateway / kgateway / Istio)
    participant EPP as Endpoint Picker<br/>(llm-d router)
    participant P as InferencePool<br/>(vLLM pods)
    C->>GW: POST /v1/chat/completions
    GW->>EPP: ext-proc: headers + body (model, prompt)
    EPP->>EPP: Score pods: prefix-cache hit, queue depth,<br/>KV utilization, LoRA loaded
    EPP-->>GW: Target pod endpoint
    GW->>P: Forward to selected vLLM pod
    P-->>C: Streamed tokens

Prefill/Decode Disaggregation

LLM inference has two phases with opposite profiles. Prefill processes the whole prompt in parallel and is compute-bound (it sets time to first token). Decode generates one token per step and is memory-bandwidth-bound (it sets inter-token latency). Running both on the same GPU makes long prefills stall everyone's decode.

Disaggregation runs separate prefill and decode worker pools and ships the KV cache from prefill to decode over NVLink or RDMA (NIXL in both llm-d and Dynamo). Each pool can then be sized, parallelized and autoscaled for its own bottleneck. The gains are workload-dependent. llm-d cites up to 70% higher tokens/s on GPT-OSS on B200 (AWS) and 10-30% on MI300X (Oracle) versus standard vLLM, per its README. It adds operational cost: two pools, a KV transfer fabric, and gang placement of the pair (Grove, KAI hierarchical PodGroups, or LeaderWorkerSet).


Full Platform Architecture Diagram

How the layers and projects on this page fit together in one cluster:

graph TB
    subgraph "Layer 1: AI Applications"
        App1["Assistants and agents"]
        App2["Recommendation Systems"]
        App3["Fraud Detection"]
        App4["Content Generation"]
    end

    subgraph "Layer 2: AI Platform Services"
        subgraph "Data Pipeline"
            DP["Data Processing<br/>(Spark, Ray Data)"]
            FE["Feature Engineering<br/>(Feast)"]
        end
        subgraph "Model Lifecycle"
            MT["Model Training<br/>(Kubeflow Trainer, Ray Train)"]
            MR["Model Registry<br/>(MLflow)"]
        end
        subgraph "Serving & Monitoring"
            IGW["Inference Gateway<br/>(Gateway API + EPP)"]
            MS["Model Serving<br/>(KServe, llm-d, Dynamo, Ray Serve)"]
            ENG["Engines<br/>(vLLM, SGLang, TensorRT-LLM)"]
            MON["Monitoring<br/>(Prometheus, DCGM Exporter)"]
        end
    end

    subgraph "Layer 3: Infrastructure"
        subgraph "Orchestration"
            K8S["Kubernetes<br/>(kube-scheduler, DRA)"]
            SCHED["Queueing and gang scheduling<br/>(Kueue, Volcano, KAI)"]
            RAY["KubeRay / Ray Cluster"]
        end
        subgraph "Compute & Storage"
            GPU["GPUs<br/>(A100, H100, B200)"]
            CPU["CPUs"]
            STORE["Object Storage<br/>(S3, MinIO)"]
        end
        subgraph "GPU Management"
            GPUOP["GPU Operator"]
            DEVPLUGIN["Device Plugin /<br/>DRA driver"]
            MIG["MIG Manager"]
        end
    end

    App1 & App2 & App3 & App4 --> IGW
    IGW --> MS --> ENG
    DP --> FE --> MT --> MR --> MS
    ENG --> MON
    MS --> RAY
    MT --> RAY
    MT --> SCHED
    RAY --> K8S
    SCHED --> K8S
    K8S --> GPU & CPU & STORE
    GPUOP --> DEVPLUGIN --> GPU
    GPUOP --> MIG --> GPU

Benchmarks and Scale Considerations

vLLM Performance Characteristics

Based on the PagedAttention paper (arXiv:2309.06180, 2023, vLLM v0.1-era measurements):

  • PagedAttention cuts KV cache waste to under 4%, versus 60-80% in the fixed-reservation allocators it measured
  • vLLM delivered 2-4x the throughput of FasterTransformer and Orca at the same latency. The paper attributes the gain to fitting more sequences in memory, with continuous batching as a prerequisite, not to batching alone
  • Memory savings translate directly to higher concurrent request capacity

Engine benchmarks age fast

Absolute numbers from 2023 no longer describe current vLLM (V1 engine, CUDA graphs, FP8/FP4 kernels). Benchmark on your own model, hardware and traffic shape before choosing an engine or GPU count.

GPU Memory Budget (Inference)

For an LLM with P parameters at B bytes per parameter:

Model weights:  P × B bytes
KV cache:       2 × layers × kv_heads × head_dim × bytes × tokens_in_flight
Overhead:       ~10-20% for framework, activations, CUDA context and graphs

Example: Llama 2 70B at FP16:

  • Model weights: 70B × 2 bytes = 140 GB
  • Minimum: 2x A100 80GB (tensor parallelism) leaves almost no KV cache
  • With KV cache headroom: 4x A100 80GB for production batch sizes
  • The same model in FP8 (about 70 GB) fits 2x H100 80GB with room for KV cache

Grouped-query attention (fewer KV heads) and FP8 KV cache (--kv-cache-dtype fp8) are the two biggest levers on KV memory in current models.

Interconnect and Topology

Inter-node bandwidth per GPU is roughly 10-30x lower than intra-node NVLink, so tensor parallelism should stay inside an NVLink domain, with pipeline or data parallelism across nodes. See Reference for per-generation figures.


Security

AI platform infrastructure introduces security concerns beyond traditional Kubernetes workloads due to the high value of GPU resources, model weights, and training data. Setup recipes are in How-to Guides. Checklists and port tables are in Reference.

Threat Model Overview

The main attack paths against an AI platform and the assets each one targets:

graph TB
    subgraph "Attack Surface"
        A1["Model Weight Theft"]
        A2["GPU Resource Hijacking<br/>(Cryptomining)"]
        A3["Training Data Exfiltration"]
        A4["Inference API Abuse<br/>(prompt injection, scraping)"]
        A5["Supply Chain<br/>(Poisoned Models, pickle payloads)"]
        A6["Multi-Tenant Isolation<br/>Bypass"]
        A7["Unauthenticated Ray Jobs API<br/>(remote code execution)"]
    end

    subgraph "Assets at Risk"
        M["Model Weights<br/>(Proprietary IP)"]
        G["GPU Compute<br/>(High cost per hour)"]
        D["Training Data<br/>(PII, proprietary)"]
        I["Inference Endpoints<br/>(Production services)"]
    end

    A1 --> M
    A2 --> G
    A3 --> D
    A4 --> I
    A5 --> M
    A6 --> G & M & D
    A7 --> G & D

Authentication and Authorization

GPU access control has two halves. Who may consume GPUs is a Kubernetes RBAC and quota problem: RBAC on pods, jobs and ResourceClaims, plus ResourceQuota (or Kueue/KAI queues) per team. Who may call a model is an application-layer problem. vLLM only offers a single static --api-key. Ray's dashboard and Jobs API historically had no authentication at all, and exposed Ray clusters have been exploited for cryptomining and data theft. Identity, per-tenant rate limits and audit therefore belong in a gateway (an AI gateway or Gateway API implementation), with NetworkPolicies ensuring nothing bypasses it, and mTLS inside the mesh.

Multi-Tenant GPU Isolation

MIG vs Time-Slicing Security Comparison

Dimension MIG Time-Slicing
Memory isolation Hardware-enforced: separate memory controllers and DRAM paths None between pods. All contexts share the GPU's framebuffer (MPS can add per-client limits)
Compute isolation Dedicated SMs, L2 cache banks, GPU engines Shared GPU with time-based access rotation
Fault isolation Per instance Shared: one fatal error can reset all contexts
Side-channel risk Low: physically separate paths Higher: shared cache and memory bus
QoS guarantees Predictable: dedicated resources Variable: depends on co-tenant behavior
Suitable for Multi-tenant production, compliance environments Development, trusted single-tenant clusters

Multi-Tenant Isolation

Time-slicing does NOT provide hardware-level isolation. Workloads from different tenants sharing a GPU via time-slicing can potentially observe side-channel information through shared L2 cache timing. For regulated or untrusted multi-tenant environments, MIG or dedicated GPU allocation is required. For hostile tenants, add sandboxed runtimes (Kata Containers with GPU passthrough) or Confidential Computing on H100 and later.

Namespace-Based GPU Isolation

Taints, tolerations and node selectors can reserve whole GPU node pools for one team. This trades utilization for a hard blast-radius boundary and is often combined with MIG for the shared pool. See How-to Guides.

Model Weight Protection

Securing Model Storage

Model weights represent significant IP investment. Protection strategies:

  1. Encrypted storage at rest: use encrypted PersistentVolumes or encrypted object storage (S3 SSE, MinIO encryption)
  2. Access control: restrict model registry access (MLflow, Hugging Face Hub) with per-team credentials
  3. Network segmentation: model downloads should traverse private networks, not the public internet
  4. Signed models: verify model integrity using checksums or signatures before loading. Prefer safetensors over pickle-based formats, which can execute code on load

GPU Cryptomining Prevention

GPU nodes are high-value targets for cryptomining. The defense is layered: admission (only approved images and workload patterns may request GPUs), quota (a compromised namespace cannot grab the fleet), and detection (sustained high utilization from pods outside approved workload labels). Detection queries are in How-to Guides.

Encryption

Data in transit and at rest matrices are in Reference. The notable AI-specific gaps: the KV cache and weights in GPU memory are plaintext unless Confidential Computing is enabled. KV cache offload tiers (CPU memory, local SSD, remote stores used by LMCache, llm-d and Dynamo KVBM) persist prompt-derived data outside the GPU and need the same protection as logs.

Compliance Considerations

Regulated environments need to prove who ran what on which GPU and which data and model versions were involved. That means audit logs for pod and ResourceClaim creation and for queue objects, lineage from training data to model registry entries, and an inventory of which tenants shared which physical GPUs (MIG layouts). The checklist is in Reference.

Sources