AI Platform Engineering — Explanation¶
Why the AI platform stack looks the way it does: how Kubernetes learns about GPUs (device plugins, then DRA), why GPUs fragment and how sharing works, why AI jobs need gang and topology-aware scheduling, how Ray and vLLM schedule work above Kubernetes, and how the LLM-aware serving layer (KServe, Gateway API Inference Extension, llm-d, NVIDIA Dynamo) routes requests by KV cache. Versions and look-up tables live in Reference. Recipes live in How-to Guides.
The Three-Layer AI Platform Architecture¶
A production AI platform consists of three distinct layers:
Layer 1: AI Applications (Business Layer)¶
The products users interact with:
- Virtual assistants, chatbots and agents
- Recommendation systems
- Fraud detection pipelines
- Content generation services
- IoT analytics
Layer 2: AI Platform Services (MLOps Layer)¶
The platform that takes raw data and turns it into production models:
| Stage | Purpose | Example Tools |
|---|---|---|
| Data Processing | Collect, clean, label, store data | Spark, Flink, Ray Data, Label Studio |
| Feature Engineering | Transform raw data into model features | Feature Stores (Feast, Tecton) |
| Model Training | Distributed GPU training | Kubeflow Trainer (TrainJob), Ray Train |
| Model Registry | Version, store, track models | MLflow, Weights & Biases |
| Deployment & Inference | Serve predictions reliably | KServe, Ray Serve, vLLM, llm-d, NVIDIA Dynamo |
| Monitoring | Track latency, drift, costs | Prometheus, Grafana, DCGM Exporter, Evidently |
Layer 3: Infrastructure (Compute Layer)¶
The foundation everything depends on:
| Category | Components |
|---|---|
| Compute | CPUs, GPUs, TPUs |
| Storage | Object storage, data lakes, feature stores |
| Orchestration | Kubernetes, networking, scheduling (Kueue, Volcano, KAI Scheduler) |
| Accelerators | NVIDIA GPUs (A100, H100/H200, B200/GB200), AMD Instinct, Google TPUs |
GPU Discovery in Kubernetes¶
Kubernetes does not natively understand GPU hardware. The classic discovery path makes GPUs visible as schedulable extended resources:
flowchart TB
HW["GPU hardware<br/>(PCIe / SXM)"] --> DRV["NVIDIA driver<br/>(kernel module)"]
DRV --> CTK["NVIDIA Container Toolkit<br/>(CDI spec)"]
CTK --> DP["NVIDIA device plugin<br/>(DaemonSet)"]
DP -->|"gRPC ListAndWatch"| KL["kubelet<br/>node.status.allocatable"]
KL --> SCH["kube-scheduler<br/>matches nvidia.com/gpu requests"]
SCH --> POD["Pod with GPU devices injected"]
Step 1: Physical GPU + Driver¶
The physical GPU is attached to a worker node. The operating system exposes it through vendor-specific drivers. For NVIDIA GPUs, the NVIDIA driver must be installed on the node before the OS can interact with the hardware.
Step 2: Kubernetes Awareness Gap¶
Even with the driver installed, Kubernetes remains unaware of the GPU. The built-in resources the scheduler counts are:
- CPU (
cpu) - Memory (
memory) - Ephemeral storage (
ephemeral-storage) - Huge pages (
hugepages-<size>)
GPUs must be explicitly registered, either through the Device Plugin API or through DRA.
Step 3: Device Plugin Registration¶
A Device Plugin is a Kubernetes extension that advertises specialized hardware to the kubelet. The NVIDIA Device Plugin runs as a DaemonSet on each GPU node and reports available GPUs.
Once registered, Kubernetes sees the resource:
Key Insight
Kubernetes does not schedule GPUs because it understands GPUs. It schedules GPUs because a Device Plugin exposes them as generic extended resources, which are opaque integer counters. The same mechanism works for TPUs, FPGAs, SmartNICs, and any other accelerator with a Device Plugin implementation. That opacity is the limitation DRA was built to remove.
Step 4: Pod GPU Requests¶
After registration, workloads request GPUs identically to CPU and memory:
apiVersion: v1
kind: Pod
metadata:
name: gpu-inference
spec:
containers:
- name: model-server
image: vllm/vllm-openai:v0.30.0
resources:
limits:
nvidia.com/gpu: 1
The scheduler matches resource requests with available resources. It has no knowledge of whether the workload is running an LLM, image generator, or training job, nor of GPU model, memory size or NVLink topology beyond what node labels say.
Device Plugin Architecture¶
The registration and allocation handshake between the device plugin, kubelet and scheduler:
sequenceDiagram
participant GPU as GPU Hardware
participant Driver as NVIDIA Driver
participant DP as NVIDIA Device Plugin
participant Kubelet as kubelet
participant Scheduler as kube-scheduler
participant Pod as Pod
GPU->>Driver: Hardware attached
Driver->>DP: GPU devices available (NVML)
DP->>Kubelet: Register via gRPC<br/>ListAndWatch(nvidia.com/gpu: 4)
Kubelet->>Scheduler: Node allocatable updated
Pod->>Scheduler: Request nvidia.com/gpu: 1
Scheduler->>Kubelet: Bind Pod to GPU node
Kubelet->>DP: Allocate(deviceID)
DP->>Pod: CDI device names / device nodes + env vars
The Device Plugin communicates with the kubelet via a gRPC interface over Unix sockets in /var/lib/kubelet/device-plugins/. The plugin must handle kubelet restarts by monitoring socket deletion and re-registering. The Device Plugin API supports:
ListAndWatch— advertises available devices and reports health changesAllocate— provisions device access for containers (device nodes, environment variables, mounts, CDI device names)- Health monitoring — marks devices as unhealthy when failures are detected. This reduces the node's allocatable count
Since Kubernetes v1.36 (beta, KEP-4680), allocatedResourcesStatus in pod status reports per-device health for both device plugins and DRA drivers.
NVIDIA GPU Operator¶
The NVIDIA GPU Operator packages the whole node stack (containerized driver, Container Toolkit, device plugin, GPU Feature Discovery, DCGM Exporter, MIG Manager) as operands of one ClusterPolicy custom resource. It removes the need for a special GPU OS image: a new node joins, NFD labels it as NVIDIA hardware, and the operator rolls the stack onto it. The component list and versions per release are in Reference.
Dynamic Resource Allocation (DRA)¶
DRA (resource.k8s.io/v1, GA since Kubernetes 1.34) replaces "count of opaque devices" with structured device descriptions the scheduler can reason about. A DRA driver publishes each node's devices and their attributes (model, memory, MIG capability, NVLink/PCIe topology) as ResourceSlice objects. Workloads create ResourceClaims, directly or via ResourceClaimTemplates, that select devices from a DeviceClass with CEL expressions. The scheduler allocates concrete devices before binding the pod. The kubelet plugin then prepares them (via CDI) at pod start.
How a GPU claim flows through DRA, from driver inventory to container start:
sequenceDiagram
participant Drv as DRA kubelet plugin<br/>(gpu.nvidia.com)
participant API as kube-apiserver
participant User as Workload owner
participant Sch as kube-scheduler<br/>(DRA plugin)
participant Kl as kubelet
Drv->>API: Publish ResourceSlice<br/>(GPUs + attributes per node)
User->>API: Create Pod + ResourceClaimTemplate<br/>(deviceClassName gpu.nvidia.com)
API->>API: Generate ResourceClaim for the Pod
Sch->>API: Read ResourceSlices, evaluate CEL selectors
Sch->>API: Write allocation to ResourceClaim status, bind Pod
Kl->>Drv: NodePrepareResources(claim)
Drv->>Kl: CDI device IDs (GPU, MIG slice, or IMEX channel)
Kl->>Kl: Start containers with devices injected
Why this matters for AI platforms:
| Device plugin limitation | DRA answer |
|---|---|
Only integer counts (nvidia.com/gpu: 1) |
Requests by attribute: "one GPU with at least 80 GiB", "an H100 or else two A100s" (prioritized list, GA in 1.36) |
| MIG layouts must be pre-sliced on the node | Partitionable devices (beta in 1.36) let the scheduler pick a partition size at scheduling time |
| Sharing needs per-plugin hacks | One ResourceClaim can be referenced by several containers or pods |
| No way to drain a bad GPU gracefully | Device taints and tolerations (GA in 1.37) |
| Migration requires rewriting every manifest | Extended resource mapping (GA in 1.37) lets a DRA driver satisfy classic nvidia.com/gpu requests |
| Multi-node NVLink is invisible | NVIDIA ComputeDomains allocate IMEX domains across GB200/GB300 NVL72 nodes |
NVIDIA donated its DRA driver to the Kubernetes project at KubeCon Europe 2026. It now lives at kubernetes-sigs/dra-driver-nvidia-gpu and ships from registry.k8s.io. Its ComputeDomain side is production-ready. GPU allocation was still flagged "not yet officially supported" in the README as of 2026-09, which is why most clusters still run the classic device plugin for plain GPUs.
Device plugins are not deprecated
DRA and device plugins coexist. The extended-resource bridge exists precisely so clusters can move node pools to DRA without changing tenant manifests. Expect a multi-year transition.
GPU Utilization and Resource Fragmentation¶
The Core Problem¶
Standard Kubernetes GPU allocation is binary. A pod requests one or more whole GPUs, and each GPU is allocated exclusively:
Pod A → 1 GPU Requested → 1 GPU Allocated (80 GB)
Actual usage: 10 GB memory, 20% compute
Waste: 70 GB memory, 80% compute
With CPUs, Kubernetes efficiently bin-packs multiple pods onto a single node. With the device plugin model, GPUs do not support this: one pod per GPU, regardless of actual utilization.
Cost Impact¶
Consider a cluster with 8x NVIDIA A100 GPUs where every workload uses only 25% of each GPU:
- Effective utilization: 2 GPUs worth of useful work
- Paid for: 8 GPUs
- Waste: 75% of GPU investment
At an illustrative purchase price of about $30,000 per A100 (2023-era list pricing; street and cloud prices vary widely and have fallen since), the six idle GPUs represent about $180,000 of stranded capacity.
GPU Sharing Strategies¶
Four techniques address GPU underutilization at the infrastructure level, plus one at the serving layer. A side-by-side table is in Reference.
Time-Slicing¶
Multiple workloads take turns using the same GPU, analogous to CPU time-sharing. The NVIDIA device plugin's native time-slicing simply advertises each GPU N times (replicas). The CUDA driver round-robins contexts with no memory limits and no fault isolation.
NVIDIA Run:ai adds scheduler-controlled time-slicing with two modes:
| Mode | Behavior | K8s Mapping |
|---|---|---|
| Strict | Each workload gets exactly its requested GPU compute fraction | gpu-compute-request = gpu-compute-limit = gpu-fraction |
| Fair | Each workload gets at least its fraction, plus unused slices from idle workloads | gpu-compute-request = gpu-fraction, gpu-compute-limit = 1.0 |
Per the Run:ai documentation, time-slicing operates on a plan/lease cycle. Default configuration:
- Lease time: 250ms (exclusive GPU access per workload)
- Granularity: 5% precision
- Plan (cycle) time: 250ms / 0.05 = 5000ms (5 seconds)
A workload requesting gpu-fraction=0.5 gets 2.5s of runtime per 5s cycle.
Trade-offs
Decreasing lease time makes time-slicing less accurate. Increasing lease time improves accuracy but reduces workload responsiveness. Context switching between workloads adds overhead.
Multi-Process Service (MPS)¶
CUDA MPS runs kernels from several processes concurrently on one GPU through a shared control daemon, instead of alternating contexts. The NVIDIA device plugin supports sharing.mps, which also sets per-client memory and thread limits. MPS gives better throughput than time-slicing for many small inference processes, but clients still share one fault domain.
Multi-Instance GPU (MIG)¶
Available on data-center GPUs from Ampere onward (A30, A100, H100/H200, GH200, B200/GB200, RTX PRO 6000 Blackwell). MIG partitions a single GPU into up to 7 isolated GPU Instances, each with dedicated:
- Streaming Multiprocessors (SMs)
- GPU engines (copy engines, decoders)
- L2 cache banks
- Memory controllers
- DRAM address busses
Example partitioning of an 80 GB A100:
Full A100 (80 GB)
├── MIG Instance 1: 10 GB (1g.10gb)
├── MIG Instance 2: 10 GB (1g.10gb)
├── MIG Instance 3: 20 GB (2g.20gb)
└── MIG Instance 4: 40 GB (4g.40gb)
The example is illustrative. Valid combinations are constrained by slice placement, so check with nvidia-smi mig -lgipp. Each instance provides hardware-level isolation: one workload cannot impact the L2 cache or DRAM bandwidth of another. This makes MIG suitable for multi-tenant environments where QoS guarantees are required. The cost is rigidity: changing a layout requires draining the GPU, which DRA partitionable devices aim to automate.
MIG supports:
- Bare-metal and containers
- GPU passthrough virtualization
- vGPU on supported hypervisors
Continuous Batching (Inference)¶
Instead of processing inference requests one-by-one, the serving engine combines multiple requests into a single GPU execution. vLLM's continuous batching dynamically adds new requests as older ones complete. The GPU then stays busy continuously instead of waiting for a full batch to form. For LLM serving this usually recovers more utilization than any infrastructure-level sharing.
AI Workload Scheduling¶
Why Standard Kubernetes Scheduling Falls Short¶
Traditional applications are loosely coupled: components can start independently and tolerate staggered scheduling. The default kube-scheduler places one pod at a time. AI training jobs have fundamentally different requirements.
Gang Scheduling¶
Distributed training jobs require all resources allocated simultaneously:
Training Job (requires 8 GPUs)
├── Worker 0: GPU 0
├── Worker 1: GPU 1
├── Worker 2: GPU 2
├── Worker 3: GPU 3
├── Worker 4: GPU 4
├── Worker 5: GPU 5
├── Worker 6: GPU 6
└── Worker 7: GPU 7
If only 6 GPUs are available, the job cannot start. Partial allocation wastes resources: workers wait indefinitely for the remaining GPUs, blocking other jobs, and two half-placed jobs can deadlock each other.
Gang scheduling rule: Either schedule all required resources together, or schedule none of them.
Until recently this required a replacement scheduler (Volcano, KAI) or Kueue's admission-level approximation. Kubernetes added a native Workload / PodGroup API (KEP-4671): alpha in 1.35, beta (v1beta1) in 1.37 with a hierarchical CompositePodGroup alpha, and stable planned for 1.38.
Topology Awareness¶
GPU placement affects training performance significantly. Same-node GPUs communicate via NVLink (900 GB/s per GPU on H100), while cross-node GPUs use network fabric (InfiniBand NDR at 400 Gb/s, about 50 GB/s, per port). Bandwidth figures are in Reference.
The two tiers of GPU connectivity a topology-aware scheduler has to respect:
graph LR
subgraph "Node A (Fast: NVLink)"
GPU0["GPU 0"]
GPU1["GPU 1"]
GPU2["GPU 2"]
GPU3["GPU 3"]
GPU0 <--> GPU1
GPU1 <--> GPU2
GPU2 <--> GPU3
end
subgraph "Node B (Fast: NVLink)"
GPU4["GPU 4"]
GPU5["GPU 5"]
GPU6["GPU 6"]
GPU7["GPU 7"]
GPU4 <--> GPU5
GPU5 <--> GPU6
GPU6 <--> GPU7
end
GPU3 <-.->|"Slower: InfiniBand / RoCE"| GPU4
A topology-aware scheduler prefers placing all GPUs on the same node when possible. If that is not possible, it falls back to nodes under the same leaf switch or rack (Kueue TAS and KAI TAS read Topology levels from node labels; Volcano uses HyperNodes). Rack-scale systems such as GB200 NVL72 add a third tier: an NVLink domain spanning 72 GPUs across nodes, exposed through DRA ComputeDomains.
Scheduling Tools Comparison¶
| Tool | Type | Key Capabilities |
|---|---|---|
| Kueue | Kubernetes SIG job queueing | Admission control, quotas with borrowing, fair sharing, TAS, MultiKueue. Keeps kube-scheduler |
| Volcano | CNCF incubating batch scheduler | Gang scheduling, queues, DRF, network-topology-aware placement. v1.14 added a sharding controller for mixed agent/batch clusters |
| KAI Scheduler | CNCF sandbox AI scheduler (from Run:ai) | Hierarchical queues, gang and hierarchical PodGroups, GPU fractions, time-based fairshare, DRA, TAS |
| Kubeflow Trainer | Training operator | v2 TrainJob + runtimes on JobSet. Integrates with Kueue, Volcano and KAI for gang/TAS. Training Operator v1 (PyTorchJob etc.) is in maintenance |
| NVIDIA GPU Operator | GPU lifecycle manager | Driver management, device plugins, DCGM metrics, MIG management |
The pattern most platforms converge on: Kueue (or KAI's queues) decides whether and when a job may run against a quota. A gang-capable scheduler decides where its pods land. A feature-by-feature matrix is in Reference.
Ray — Distributed Compute Framework¶
The Two-Layer Model¶
Ray schedules computation inside pods that Kubernetes has already placed:
graph TB
subgraph "Application Layer"
User["Python program"]
RayDriver["Ray Driver<br/>(ray.init)"]
end
subgraph "Ray Layer (Computation Scheduling)"
RayHead["Ray Head Node<br/>(GCS, Autoscaler, Dashboard)"]
RayWorker1["Ray Worker 1<br/>(raylet + object store)"]
RayWorker2["Ray Worker 2<br/>(raylet + object store)"]
RayWorker3["Ray Worker 3<br/>(raylet + object store)"]
end
subgraph "Kubernetes Layer (Infrastructure Scheduling)"
KubeRay["KubeRay operator<br/>(RayCluster / RayJob / RayService)"]
K8s["kube-scheduler"]
Pod1["Pod (Head)"]
Pod2["Pod (Worker)"]
Pod3["Pod (Worker)"]
Pod4["Pod (Worker)"]
end
User --> RayDriver
RayDriver --> RayHead
RayHead --> RayWorker1
RayHead --> RayWorker2
RayHead --> RayWorker3
KubeRay --> K8s
K8s --> Pod1
K8s --> Pod2
K8s --> Pod3
K8s --> Pod4
Pod1 -.- RayHead
Pod2 -.- RayWorker1
Pod3 -.- RayWorker2
Pod4 -.- RayWorker3
Kubernetes schedules infrastructure (which node should this pod run on?). Ray schedules computation (which worker executes which task? how are results collected?).
Ray Architecture¶
| Component | Role |
|---|---|
| Head Node | Runs GCS (Global Control Service), autoscaler, Ray dashboard. Also schedules tasks like worker nodes unless its num-gpus/num-cpus are set to 0. |
| Worker Nodes | Execute Ray tasks and actors. Each runs a raylet and participates in the distributed object store. |
| Autoscaler | Scales worker nodes based on task/actor resource requests (not CPU/memory metrics). |
| GCS | Central metadata store for cluster state, actor locations, and resource availability. |
Tasks vs Actors¶
| Dimension | Tasks | Actors |
|---|---|---|
| State | Stateless | Stateful |
| Lifecycle | Run once, return result | Long-lived, handle multiple requests |
| Use case | Data processing, hyperparameter search | Model serving, stateful computation |
| Invocation | function.remote(args) |
actor.method.remote(args) |
Tasks enable embarrassingly parallel workloads (data processing, hyperparameter tuning). Actors enable stateful services (model serving, game environments, RL training).
Ray Libraries¶
| Library | Purpose |
|---|---|
| Ray Train | Distributed training (PyTorch, TensorFlow, XGBoost) |
| Ray Tune | Hyperparameter optimization |
| Ray Serve | Model serving and composition, including Ray Serve LLM on vLLM |
| Ray Data | Distributed data processing, including batch LLM inference |
| Ray RLlib | Reinforcement learning |
Ray on Kubernetes (KubeRay)¶
KubeRay provides Kubernetes CRDs for managing Ray clusters:
RayCluster— manages head and worker podsRayJob— creates a cluster (or uses an existing one), submits an entrypoint, and can tear the cluster down afterwardsRayService— manages Ray Serve deployments with zero-downtime upgrades (incremental upgrades since v1.5)
With enableInTreeAutoscaling: true, KubeRay injects the Ray autoscaler as a sidecar container in the head pod. It scales worker pods between minReplicas and maxReplicas based on pending Ray task/actor resource demands. KubeRay v1.5 added Ray token authentication support. KubeRay integrates with Kueue, Volcano and KAI Scheduler for gang admission of whole Ray clusters.
vLLM — Inference Engine Architecture¶
Why Naive Model Serving Fails at Scale¶
A naive inference server processes requests sequentially or in static batches:
Request 1 → GPU → Response 1
Request 2 → GPU → Response 2 (waits for Request 1)
Request 3 → GPU → Response 3 (waits for Request 2)
With 100 concurrent users, GPU utilization stays low because the GPU cannot exploit its parallel architecture. LLM decoding is memory-bandwidth-bound: every generated token re-reads all model weights, so throughput comes from batching many sequences into each step. This is the GPU utilization problem applied to inference serving.
PagedAttention¶
The key innovation in vLLM. Traditional serving systems allocate GPU memory for the KV cache in large, contiguous chunks. For variable-length sequences, this causes:
- Internal fragmentation — allocated blocks larger than needed
- External fragmentation — free memory scattered in unusable small chunks
- Reservation waste — memory reserved for maximum sequence length even for short sequences
PagedAttention treats the KV cache like virtual memory. Memory is managed in fixed-size blocks (pages) that need not be contiguous, and a block table maps each sequence's logical blocks to physical ones:
Traditional KV Cache:
[████████████░░░░░░░░] Request 1 (wasted space)
[████████░░░░░░░░░░░░] Request 2 (wasted space)
[░░░░░░░░░░░░░░░░░░░░] Free (fragmented)
PagedAttention:
[████][████][████][██] Request 1 (pages, no waste)
[████][████][██] Request 2 (pages, no waste)
[████][████] Free (reusable pages)
Results reported in the PagedAttention paper (arXiv:2309.06180):
- Existing systems wasted 60-80% of KV cache memory. vLLM keeps waste under 4% (only the last partial block)
- 2-4x higher throughput than FasterTransformer and Orca at the same latency
- Block-level sharing enables copy-on-write for parallel sampling and beam search, and later automatic prefix caching
Continuous Batching¶
Traditional batching waits for a full batch before processing:
Traditional: Wait → Process Batch → Wait → Process Batch
Continuous: Process ─── Process ─── Process ─── Process
(new requests added as old ones complete)
vLLM's continuous (iteration-level) batching adds new requests to the running batch as existing requests finish generating tokens. The GPU stays busy continuously. The V1 engine (the only engine in current releases) also uses chunked prefill: long prompts are split so prefill chunks and decode tokens share each step under a --max-num-batched-tokens budget.
vLLM in the Platform Stack¶
Users
|
vLLM (inference optimization: PagedAttention + continuous batching)
|
Model Weights (loaded into GPU memory)
|
GPU (compute)
In a Kubernetes deployment:
Users
|
Gateway (Gateway API + Inference Extension / AI gateway)
|
vLLM Pods (plain Deployment, KServe LLMInferenceService, llm-d, Dynamo, or Ray Serve)
|
Kubernetes (scheduling, scaling, health checks)
|
GPU Nodes (NVIDIA Device Plugin or DRA driver, GPU Operator)
vLLM exposes an OpenAI-compatible API. It is a drop-in replacement for existing applications that use the OpenAI API format.
LLM-Aware Serving Layer¶
Once an LLM runs on more than a handful of replicas, the load balancer becomes a performance component. Round-robin ignores which replica already holds a request's prefix in KV cache, how deep each replica's queue is, and which LoRA adapters are loaded. Since 2025 a Kubernetes-native layer has formed to fix that.
| Project | What it adds | Relationship |
|---|---|---|
| Gateway API Inference Extension (GAIE) | InferencePool API (GA, inference.networking.k8s.io/v1) and the ext-proc Endpoint Picker (EPP) protocol that turns Envoy Gateway, kgateway, Istio or GKE Gateway into an "inference gateway" |
Standard API. As of 2026 the full EPP, InferenceObjective and body-based router moved to llm-d/llm-d-router. GAIE keeps InferencePool, a lightweight reference EPP and conformance tests |
| llm-d | Well-lit paths on top of vLLM: prefix-cache and load-aware routing, tiered KV cache offload, prefill/decode disaggregation, wide expert parallelism | CNCF sandbox (2026-03), founded by Red Hat, Google Cloud, IBM Research, CoreWeave and NVIDIA |
| KServe | LLMInferenceService CRD (router, prefill, worker, parallelism, KV offload) alongside the classic InferenceService |
CNCF incubating. Uses llm-d scheduling and GAIE resources under the hood |
| NVIDIA Dynamo | Engine-agnostic (vLLM, SGLang, TensorRT-LLM) frontend, KV-aware router, SLA-based Planner autoscaler, KV Block Manager (GPU → CPU → SSD → remote), NIXL transfers, Grove gang-scheduling operator | Apache-2.0. Can run its own frontend or plug into GAIE as an EPP |
| vLLM production-stack | Helm chart with a session/prefix-aware router and LMCache offload | Lighter-weight reference stack from the vLLM project |
The request path through an inference gateway with an Endpoint Picker, which is the common shape across GAIE, llm-d and KServe:
sequenceDiagram
participant C as Client (OpenAI API)
participant GW as Gateway<br/>(Envoy Gateway / kgateway / Istio)
participant EPP as Endpoint Picker<br/>(llm-d router)
participant P as InferencePool<br/>(vLLM pods)
C->>GW: POST /v1/chat/completions
GW->>EPP: ext-proc: headers + body (model, prompt)
EPP->>EPP: Score pods: prefix-cache hit, queue depth,<br/>KV utilization, LoRA loaded
EPP-->>GW: Target pod endpoint
GW->>P: Forward to selected vLLM pod
P-->>C: Streamed tokens
Prefill/Decode Disaggregation¶
LLM inference has two phases with opposite profiles. Prefill processes the whole prompt in parallel and is compute-bound (it sets time to first token). Decode generates one token per step and is memory-bandwidth-bound (it sets inter-token latency). Running both on the same GPU makes long prefills stall everyone's decode.
Disaggregation runs separate prefill and decode worker pools and ships the KV cache from prefill to decode over NVLink or RDMA (NIXL in both llm-d and Dynamo). Each pool can then be sized, parallelized and autoscaled for its own bottleneck. The gains are workload-dependent. llm-d cites up to 70% higher tokens/s on GPT-OSS on B200 (AWS) and 10-30% on MI300X (Oracle) versus standard vLLM, per its README. It adds operational cost: two pools, a KV transfer fabric, and gang placement of the pair (Grove, KAI hierarchical PodGroups, or LeaderWorkerSet).
Full Platform Architecture Diagram¶
How the layers and projects on this page fit together in one cluster:
graph TB
subgraph "Layer 1: AI Applications"
App1["Assistants and agents"]
App2["Recommendation Systems"]
App3["Fraud Detection"]
App4["Content Generation"]
end
subgraph "Layer 2: AI Platform Services"
subgraph "Data Pipeline"
DP["Data Processing<br/>(Spark, Ray Data)"]
FE["Feature Engineering<br/>(Feast)"]
end
subgraph "Model Lifecycle"
MT["Model Training<br/>(Kubeflow Trainer, Ray Train)"]
MR["Model Registry<br/>(MLflow)"]
end
subgraph "Serving & Monitoring"
IGW["Inference Gateway<br/>(Gateway API + EPP)"]
MS["Model Serving<br/>(KServe, llm-d, Dynamo, Ray Serve)"]
ENG["Engines<br/>(vLLM, SGLang, TensorRT-LLM)"]
MON["Monitoring<br/>(Prometheus, DCGM Exporter)"]
end
end
subgraph "Layer 3: Infrastructure"
subgraph "Orchestration"
K8S["Kubernetes<br/>(kube-scheduler, DRA)"]
SCHED["Queueing and gang scheduling<br/>(Kueue, Volcano, KAI)"]
RAY["KubeRay / Ray Cluster"]
end
subgraph "Compute & Storage"
GPU["GPUs<br/>(A100, H100, B200)"]
CPU["CPUs"]
STORE["Object Storage<br/>(S3, MinIO)"]
end
subgraph "GPU Management"
GPUOP["GPU Operator"]
DEVPLUGIN["Device Plugin /<br/>DRA driver"]
MIG["MIG Manager"]
end
end
App1 & App2 & App3 & App4 --> IGW
IGW --> MS --> ENG
DP --> FE --> MT --> MR --> MS
ENG --> MON
MS --> RAY
MT --> RAY
MT --> SCHED
RAY --> K8S
SCHED --> K8S
K8S --> GPU & CPU & STORE
GPUOP --> DEVPLUGIN --> GPU
GPUOP --> MIG --> GPU
Benchmarks and Scale Considerations¶
vLLM Performance Characteristics¶
Based on the PagedAttention paper (arXiv:2309.06180, 2023, vLLM v0.1-era measurements):
- PagedAttention cuts KV cache waste to under 4%, versus 60-80% in the fixed-reservation allocators it measured
- vLLM delivered 2-4x the throughput of FasterTransformer and Orca at the same latency. The paper attributes the gain to fitting more sequences in memory, with continuous batching as a prerequisite, not to batching alone
- Memory savings translate directly to higher concurrent request capacity
Engine benchmarks age fast
Absolute numbers from 2023 no longer describe current vLLM (V1 engine, CUDA graphs, FP8/FP4 kernels). Benchmark on your own model, hardware and traffic shape before choosing an engine or GPU count.
GPU Memory Budget (Inference)¶
For an LLM with P parameters at B bytes per parameter:
Model weights: P × B bytes
KV cache: 2 × layers × kv_heads × head_dim × bytes × tokens_in_flight
Overhead: ~10-20% for framework, activations, CUDA context and graphs
Example: Llama 2 70B at FP16:
- Model weights: 70B × 2 bytes = 140 GB
- Minimum: 2x A100 80GB (tensor parallelism) leaves almost no KV cache
- With KV cache headroom: 4x A100 80GB for production batch sizes
- The same model in FP8 (about 70 GB) fits 2x H100 80GB with room for KV cache
Grouped-query attention (fewer KV heads) and FP8 KV cache (--kv-cache-dtype fp8) are the two biggest levers on KV memory in current models.
Interconnect and Topology¶
Inter-node bandwidth per GPU is roughly 10-30x lower than intra-node NVLink, so tensor parallelism should stay inside an NVLink domain, with pipeline or data parallelism across nodes. See Reference for per-generation figures.
Security¶
AI platform infrastructure introduces security concerns beyond traditional Kubernetes workloads due to the high value of GPU resources, model weights, and training data. Setup recipes are in How-to Guides. Checklists and port tables are in Reference.
Threat Model Overview¶
The main attack paths against an AI platform and the assets each one targets:
graph TB
subgraph "Attack Surface"
A1["Model Weight Theft"]
A2["GPU Resource Hijacking<br/>(Cryptomining)"]
A3["Training Data Exfiltration"]
A4["Inference API Abuse<br/>(prompt injection, scraping)"]
A5["Supply Chain<br/>(Poisoned Models, pickle payloads)"]
A6["Multi-Tenant Isolation<br/>Bypass"]
A7["Unauthenticated Ray Jobs API<br/>(remote code execution)"]
end
subgraph "Assets at Risk"
M["Model Weights<br/>(Proprietary IP)"]
G["GPU Compute<br/>(High cost per hour)"]
D["Training Data<br/>(PII, proprietary)"]
I["Inference Endpoints<br/>(Production services)"]
end
A1 --> M
A2 --> G
A3 --> D
A4 --> I
A5 --> M
A6 --> G & M & D
A7 --> G & D
Authentication and Authorization¶
GPU access control has two halves. Who may consume GPUs is a Kubernetes RBAC and quota problem: RBAC on pods, jobs and ResourceClaims, plus ResourceQuota (or Kueue/KAI queues) per team. Who may call a model is an application-layer problem. vLLM only offers a single static --api-key. Ray's dashboard and Jobs API historically had no authentication at all, and exposed Ray clusters have been exploited for cryptomining and data theft. Identity, per-tenant rate limits and audit therefore belong in a gateway (an AI gateway or Gateway API implementation), with NetworkPolicies ensuring nothing bypasses it, and mTLS inside the mesh.
Multi-Tenant GPU Isolation¶
MIG vs Time-Slicing Security Comparison¶
| Dimension | MIG | Time-Slicing |
|---|---|---|
| Memory isolation | Hardware-enforced: separate memory controllers and DRAM paths | None between pods. All contexts share the GPU's framebuffer (MPS can add per-client limits) |
| Compute isolation | Dedicated SMs, L2 cache banks, GPU engines | Shared GPU with time-based access rotation |
| Fault isolation | Per instance | Shared: one fatal error can reset all contexts |
| Side-channel risk | Low: physically separate paths | Higher: shared cache and memory bus |
| QoS guarantees | Predictable: dedicated resources | Variable: depends on co-tenant behavior |
| Suitable for | Multi-tenant production, compliance environments | Development, trusted single-tenant clusters |
Multi-Tenant Isolation
Time-slicing does NOT provide hardware-level isolation. Workloads from different tenants sharing a GPU via time-slicing can potentially observe side-channel information through shared L2 cache timing. For regulated or untrusted multi-tenant environments, MIG or dedicated GPU allocation is required. For hostile tenants, add sandboxed runtimes (Kata Containers with GPU passthrough) or Confidential Computing on H100 and later.
Namespace-Based GPU Isolation¶
Taints, tolerations and node selectors can reserve whole GPU node pools for one team. This trades utilization for a hard blast-radius boundary and is often combined with MIG for the shared pool. See How-to Guides.
Model Weight Protection¶
Securing Model Storage¶
Model weights represent significant IP investment. Protection strategies:
- Encrypted storage at rest: use encrypted PersistentVolumes or encrypted object storage (S3 SSE, MinIO encryption)
- Access control: restrict model registry access (MLflow, Hugging Face Hub) with per-team credentials
- Network segmentation: model downloads should traverse private networks, not the public internet
- Signed models: verify model integrity using checksums or signatures before loading. Prefer
safetensorsover pickle-based formats, which can execute code on load
GPU Cryptomining Prevention¶
GPU nodes are high-value targets for cryptomining. The defense is layered: admission (only approved images and workload patterns may request GPUs), quota (a compromised namespace cannot grab the fleet), and detection (sustained high utilization from pods outside approved workload labels). Detection queries are in How-to Guides.
Encryption¶
Data in transit and at rest matrices are in Reference. The notable AI-specific gaps: the KV cache and weights in GPU memory are plaintext unless Confidential Computing is enabled. KV cache offload tiers (CPU memory, local SSD, remote stores used by LMCache, llm-d and Dynamo KVBM) persist prompt-derived data outside the GPU and need the same protection as logs.
Compliance Considerations¶
Regulated environments need to prove who ran what on which GPU and which data and model versions were involved. That means audit logs for pod and ResourceClaim creation and for queue objects, lineage from training data to model registry entries, and an inventory of which tenants shared which physical GPUs (MIG layouts). The checklist is in Reference.
Sources¶
- Kubernetes Device Plugins
- Kubernetes Dynamic Resource Allocation
- Kubernetes v1.36 DRA updates and v1.37 DRA updates
- DRA Driver for NVIDIA GPUs and NVIDIA KubeCon 2026 announcement
- NVIDIA MIG User Guide
- Run:ai GPU Time-Slicing
- KAI Scheduler, Kueue, Volcano, Kubeflow Trainer
- Ray Cluster Key Concepts and Ray on Kubernetes
- PagedAttention Paper (arXiv:2309.06180)
- Gateway API Inference Extension, llm-d, KServe, NVIDIA Dynamo