AI Platform Engineering — How-to Guides¶
Task recipes for running GPU workloads on Kubernetes: prepare GPU nodes, allocate and share GPUs, schedule batch jobs, run Ray and vLLM, secure the platform, and troubleshoot. Version pins match the Reference snapshot (2026-09-25). The reasoning behind each tool is in Explanation.
Pin versions
Every command below pins a release that existed on 2026-09-25. Replace pins with the current release when you run them, and avoid :latest images in production.
GPU Node Setup¶
Prerequisites¶
Before Kubernetes can schedule GPU workloads, each GPU node requires:
- NVIDIA Driver — kernel module for GPU hardware access
- NVIDIA Container Toolkit — enables container runtimes (containerd, CRI-O) to access GPUs through CDI
- NVIDIA Device Plugin (classic) or DRA Driver for NVIDIA GPUs (DRA) — advertises GPUs to Kubernetes
- Optionally: NVIDIA GPU Operator — automates all of the above
Manual GPU Node Setup¶
# Verify GPU hardware is detected
lspci | grep -i nvidia
# Check NVIDIA driver installation
nvidia-smi
# Verify NVIDIA Container Toolkit
nvidia-container-cli info
# Deploy NVIDIA Device Plugin as DaemonSet (kube-system namespace)
kubectl apply -f https://raw.githubusercontent.com/NVIDIA/k8s-device-plugin/v0.20.1/deployments/static/nvidia-device-plugin.yml
# Verify GPU resources are advertised
kubectl describe node <gpu-node> | grep nvidia.com/gpu
Helm is the supported install path
The static manifest is for quick tests. For production use the device plugin Helm chart (https://nvidia.github.io/k8s-device-plugin) or, better, the GPU Operator, which bundles the same plugin version.
GPU Operator Deployment (Recommended)¶
The GPU Operator automates the full lifecycle (driver, toolkit, device plugin, GFD, DCGM Exporter, MIG Manager). Component versions per release are in Reference.
# Add NVIDIA Helm repo
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
# Install GPU Operator (driver, toolkit, device plugin, DCGM exporter, MIG manager are on by default)
helm install gpu-operator nvidia/gpu-operator \
--namespace gpu-operator \
--create-namespace \
--version v26.7.1 \
--wait
# If the node image already ships drivers / toolkit (e.g. some managed node pools), disable them:
# --set driver.enabled=false --set toolkit.enabled=false
# Verify all components are running
kubectl get pods -n gpu-operator
# Check GPU resources on nodes
kubectl get nodes -o json | jq '.items[].status.allocatable | select(.["nvidia.com/gpu"])'
Verify GPU Scheduling¶
# Run a test GPU pod
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: gpu-test
spec:
restartPolicy: OnFailure
containers:
- name: cuda-test
image: nvidia/cuda:12.4.1-base-ubuntu22.04
command: ["nvidia-smi"]
resources:
limits:
nvidia.com/gpu: 1
EOF
# Check pod logs
kubectl logs gpu-test
Allocate GPUs with DRA¶
Dynamic Resource Allocation is GA since Kubernetes 1.34 (resource.k8s.io/v1). The DRA Driver for NVIDIA GPUs publishes GPUs as ResourceSlices and pods request them through ResourceClaims. See Explanation for how it differs from device plugins.
GPU allocation through the NVIDIA DRA driver is still maturing
As of 2026-09 the driver README states that GPU allocation features "are not yet officially supported" and the GPU kubelet plugin is disabled by default. The install docs enable it with gpuResourcesEnabledOverride=true. ComputeDomains (Multi-Node NVLink on GB200/GB300) are the supported path. Do not run the classic device plugin and the DRA GPU plugin for the same GPUs. If the GPU Operator is installed, follow its DRA install guide instead.
# Kubernetes 1.34+ with DRA enabled (default from 1.34)
helm install dra-driver-nvidia-gpu \
oci://registry.k8s.io/dra-driver-nvidia/charts/dra-driver-nvidia-gpu \
--version 0.5.0 \
--namespace dra-driver-nvidia-gpu --create-namespace \
--set gpuResourcesEnabledOverride=true
# On GKE add: --set nvidiaDriverRoot=/home/kubernetes/bin/nvidia
# Verify: DeviceClasses gpu.nvidia.com, mig.nvidia.com, compute-domain-*.nvidia.com
kubectl get deviceclass
kubectl get resourceslice -o wide
Request one GPU shared by two containers in the same pod (from the driver's quickstart gpu-test2.yaml):
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
namespace: gpu-test2
name: single-gpu
spec:
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu.nvidia.com
---
apiVersion: v1
kind: Pod
metadata:
namespace: gpu-test2
name: pod
spec:
containers:
- name: ctr0
image: ubuntu:22.04
command: ["bash", "-c"]
args: ["nvidia-smi -L; trap 'exit 0' TERM; sleep 9999 & wait"]
resources:
claims:
- name: shared-gpu
- name: ctr1
image: ubuntu:22.04
command: ["bash", "-c"]
args: ["nvidia-smi -L; trap 'exit 0' TERM; sleep 9999 & wait"]
resources:
claims:
- name: shared-gpu
resourceClaims:
- name: shared-gpu
resourceClaimTemplateName: single-gpu
MIG Configuration¶
MIG profile tables (memory, SMs, max instances, profile IDs) are in Reference.
Enable and Configure MIG¶
# Enable MIG mode (requires GPU reset; stop all GPU processes first)
sudo nvidia-smi -i 0 -mig 1
# Reboot or reset GPU
sudo nvidia-smi -i 0 --gpu-reset
# List available MIG profiles (IDs are GPU-model specific)
nvidia-smi mig -i 0 -lgip
# Create MIG instances (IDs below are for A100 80GB)
sudo nvidia-smi mig -i 0 -cgi 19,19,14 -C
# Creates: 2x 1g.10gb + 1x 2g.20gb
# Verify MIG instances
nvidia-smi mig -i 0 -lgi
# List compute instances
nvidia-smi mig -i 0 -lci
# Destroy all MIG instances
sudo nvidia-smi mig -i 0 -dci
sudo nvidia-smi mig -i 0 -dgi
MIG with the GPU Operator¶
With the GPU Operator, do not run nvidia-smi mig by hand. Label the node and MIG Manager applies a named layout from its ConfigMap:
# Apply a uniform layout (7x 1g.10gb on A100/H100 80GB)
kubectl label nodes <gpu-node> nvidia.com/mig.config=all-1g.10gb --overwrite
# Watch the MIG Manager state
kubectl get node <gpu-node> -o jsonpath='{.metadata.labels.nvidia\.com/mig\.config\.state}'
The chart value mig.strategy decides how instances appear: single (default) exposes them as nvidia.com/gpu; mixed exposes one resource per profile.
MIG with Kubernetes¶
With mig.strategy=mixed, the NVIDIA Device Plugin advertises each MIG profile as a separate resource:
# Node capacity with MIG (mixed strategy)
# nvidia.com/mig-1g.10gb: 2
# nvidia.com/mig-2g.20gb: 1
# Pod requesting a specific MIG profile
apiVersion: v1
kind: Pod
metadata:
name: mig-workload
spec:
containers:
- name: inference
image: vllm/vllm-openai:v0.30.0
resources:
limits:
nvidia.com/mig-1g.10gb: 1
GPU Time-Slicing Configuration¶
No isolation
Time-slicing shares one GPU's memory and fault domain between all pods. A single out-of-memory or XID error can affect every pod on the GPU. Use it for dev/test, not for untrusted tenants (see Explanation).
NVIDIA Native Time-Slicing (GPU Operator)¶
# ConfigMap for NVIDIA Device Plugin time-slicing
apiVersion: v1
kind: ConfigMap
metadata:
name: time-slicing-config
namespace: gpu-operator
data:
any: |-
version: v1
flags:
migStrategy: none
sharing:
timeSlicing:
renameByDefault: false
failRequestsGreaterThanOne: false
resources:
- name: nvidia.com/gpu
replicas: 4 # Allow 4 pods per GPU
# Point the operator-managed device plugin at the ConfigMap, using key "any" as the default
kubectl patch clusterpolicies.nvidia.com/cluster-policy \
-n gpu-operator --type merge \
-p '{"spec": {"devicePlugin": {"config": {"name": "time-slicing-config", "default": "any"}}}}'
# After the device plugin restarts, allocatable nvidia.com/gpu is multiplied by the replica count
kubectl describe node <gpu-node> | grep -E "nvidia.com/gpu|gpu.replicas"
Set renameByDefault: true to advertise the shared resource as nvidia.com/gpu.shared, so workloads must opt in explicitly.
Run:ai Time-Slicing Modes¶
Run:ai (NVIDIA's commercial scheduler) has its own strict / fair time-slicing modes. They are configured in the Run:ai cluster config, not through the GPU Operator chart:
kubectl patch -n runai runaiconfigs.run.ai/runai \
--type='merge' \
--patch '{"spec":{"global":{"core":{"timeSlicing":{"mode": "fair"}}}}}'
The open-source KAI Scheduler provides GPU fractions (memory-based sharing) without a Run:ai license. See the KAI GPU sharing docs.
vLLM Deployment¶
Standalone vLLM Server¶
# Install vLLM (Linux, CUDA)
pip install vllm==0.30.0
# Start vLLM server with a model
vllm serve Qwen/Qwen3-8B \
--host 0.0.0.0 \
--port 8000 \
--tensor-parallel-size 1 \
--gpu-memory-utilization 0.9 \
--max-model-len 8192 \
--api-key "$VLLM_API_KEY"
# Test with curl (OpenAI-compatible API)
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $VLLM_API_KEY" \
-d '{
"model": "Qwen/Qwen3-8B",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 128
}'
vLLM on Kubernetes¶
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-server
spec:
replicas: 1
selector:
matchLabels:
app: vllm
template:
metadata:
labels:
app: vllm
spec:
containers:
- name: vllm
image: vllm/vllm-openai:v0.30.0
args:
- "--model"
- "Qwen/Qwen3-8B"
- "--tensor-parallel-size"
- "1"
- "--gpu-memory-utilization"
- "0.9"
- "--max-model-len"
- "8192"
ports:
- containerPort: 8000
readinessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 60
periodSeconds: 10
resources:
limits:
nvidia.com/gpu: 1
memory: 32Gi
requests:
memory: 16Gi
env:
- name: HF_TOKEN
valueFrom:
secretKeyRef:
name: hf-token
key: token
volumeMounts:
- name: dshm # shared memory for NCCL when tensor-parallel-size > 1
mountPath: /dev/shm
volumes:
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 8Gi
---
apiVersion: v1
kind: Service
metadata:
name: vllm-service
spec:
selector:
app: vllm
ports:
- port: 8000
targetPort: 8000
type: ClusterIP
Engine flags and their current defaults are in Reference.
vLLM Production Stack (Helm)¶
The vLLM project's reference stack adds a request router (session and prefix-aware), LMCache KV offloading, and Grafana dashboards:
helm repo add vllm https://vllm-project.github.io/production-stack
# values file from the repo: tutorials/assets/values-01-minimal-example.yaml
helm install vllm vllm/vllm-stack -f values-01-minimal-example.yaml
Serve an LLM with KServe LLMInferenceService¶
KServe 0.20 adds the LLMInferenceService CRD, which wires vLLM model servers to a Gateway API Inference Extension router (llm-d scheduler). Requires KServe with the LLM controller, Gateway API CRDs, and an ext-proc capable gateway (Envoy Gateway, kgateway, GKE Gateway, Istio). Example adapted from the KServe sample docs/samples/llmisvc/single-node-gpu/:
apiVersion: serving.kserve.io/v1alpha1
kind: LLMInferenceService
metadata:
name: qwen2-7b-instruct-single
spec:
model:
uri: hf://Qwen/Qwen2.5-7B-Instruct
name: Qwen/Qwen2.5-7B-Instruct
replicas: 3
router:
scheduler: { } # llm-d / EPP scheduler with default load- and prefix-aware scoring
route: { } # managed HTTPRoute
gateway: { } # attach to the default Gateway
template:
containers:
- name: main
resources:
limits:
nvidia.com/gpu: "1"
memory: 32Gi
requests:
nvidia.com/gpu: "1"
memory: 16Gi
The CRD also has prefill, worker, parallelism and kvCacheOffloading sections for disaggregated and multi-node serving. v1alpha2 is the storage version. v1alpha1 is still served.
Ray Cluster Deployment¶
Ray on Kubernetes (KubeRay)¶
# Install KubeRay operator
helm repo add kuberay https://ray-project.github.io/kuberay-helm/
helm repo update
helm install kuberay-operator kuberay/kuberay-operator \
--version 1.7.1 \
--namespace ray-system \
--create-namespace
# Deploy a Ray cluster with GPU workers and in-tree autoscaling
cat <<EOF | kubectl apply -f -
apiVersion: ray.io/v1
kind: RayCluster
metadata:
name: gpu-cluster
spec:
rayVersion: "2.57.0" # must match the image
enableInTreeAutoscaling: true # adds the autoscaler sidecar to the head pod
headGroupSpec:
rayStartParams:
num-gpus: "0" # keep GPU tasks off the head
template:
spec:
containers:
- name: ray-head
image: rayproject/ray:2.57.0-py311-gpu
ports:
- containerPort: 6379 # GCS
- containerPort: 8265 # Dashboard / Jobs API
- containerPort: 10001 # Ray Client
resources:
limits:
cpu: "4"
memory: "8Gi"
workerGroupSpecs:
- replicas: 2
minReplicas: 1
maxReplicas: 4
groupName: gpu-workers
rayStartParams: {}
template:
spec:
containers:
- name: ray-worker
image: rayproject/ray:2.57.0-py311-gpu
resources:
limits:
cpu: "4"
memory: "16Gi"
nvidia.com/gpu: 1
EOF
# Check Ray cluster status
kubectl get rayclusters
kubectl get pods -l ray.io/cluster=gpu-cluster
rayproject/ray-ml images are discontinued
Ray stopped publishing rayproject/ray-ml after 2.30.0, and the latest tags have not moved since. Use rayproject/ray:<version>-py3XX-gpu and install ML libraries in your own image. Source: Ray installation docs.
Ray Job Submission¶
# Forward the dashboard / Jobs API (service name is <cluster>-head-svc)
kubectl port-forward svc/gpu-cluster-head-svc 8265:8265
# Submit a Ray job
ray job submit \
--address http://localhost:8265 \
--working-dir . \
-- python train.py
# Check job status
ray job status <job-id> --address http://localhost:8265
# View job logs
ray job logs <job-id> --address http://localhost:8265
For CI-style runs, prefer a RayJob CR, which creates a cluster, runs the entrypoint and tears it down. KubeRay also ships a kubectl ray plugin (beta since v1.3).
Basic Ray Task Example¶
import ray
ray.init()
@ray.remote(num_gpus=1)
def train_on_chunk(data_chunk):
import torch
device = torch.device("cuda")
# Process data chunk on GPU
return len(data_chunk)
# Distribute work across GPUs
data = list(range(10_000))
chunks = [data[i:i+1000] for i in range(0, len(data), 1000)]
futures = [train_on_chunk.remote(chunk) for chunk in chunks]
results = ray.get(futures)
Batch Scheduling with Volcano¶
Install Volcano¶
# Install a released chart (recommended)
helm repo add volcano-sh https://volcano-sh.github.io/helm-charts
helm install volcano volcano-sh/volcano -n volcano-system --create-namespace
# Or the development manifest from master (not for production)
# kubectl apply -f https://raw.githubusercontent.com/volcano-sh/volcano/master/installer/volcano-development.yaml
# Verify installation
kubectl get pods -n volcano-system
Gang-Scheduled Training Job¶
The pytorch job plugin injects MASTER_ADDR, MASTER_PORT, WORLD_SIZE and RANK into every pod, and minAvailable makes Volcano place all 4 pods at once or none.
apiVersion: batch.volcano.sh/v1alpha1
kind: Job
metadata:
name: distributed-training
spec:
minAvailable: 4 # Gang scheduling: master + 3 workers required
schedulerName: volcano
queue: default
plugins:
pytorch: ["--master=master", "--worker=worker", "--port=23456"]
policies:
- event: PodEvicted
action: RestartJob
tasks:
- replicas: 1
name: master
template:
spec:
restartPolicy: OnFailure
containers:
- name: master
image: pytorch/pytorch:2.3.0-cuda12.1-cudnn8-runtime
command: ["sh", "-c"]
args:
- torchrun --nnodes=$WORLD_SIZE --node_rank=$RANK --nproc_per_node=1
--master_addr=$MASTER_ADDR --master_port=$MASTER_PORT train.py
resources:
limits:
nvidia.com/gpu: 1
- replicas: 3
name: worker
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: pytorch/pytorch:2.3.0-cuda12.1-cudnn8-runtime
command: ["sh", "-c"]
args:
- torchrun --nnodes=$WORLD_SIZE --node_rank=$RANK --nproc_per_node=1
--master_addr=$MASTER_ADDR --master_port=$MASTER_PORT train.py
resources:
limits:
nvidia.com/gpu: 1
The generic env plugin exposes the per-task index as VC_TASK_INDEX (and legacy VK_TASK_INDEX) if you need it instead.
Job Queueing with Kueue¶
Install Kueue¶
# Install Kueue (supported on Kubernetes 1.34+)
kubectl apply --server-side -f https://github.com/kubernetes-sigs/kueue/releases/download/v0.19.6/manifests.yaml
# Verify
kubectl get pods -n kueue-system
Configure Resource Quotas¶
Kueue's API is v1beta2. v1beta1 support was discontinued in v0.17, so older examples must be migrated.
# ResourceFlavor describing the GPU node pool
apiVersion: kueue.x-k8s.io/v1beta2
kind: ResourceFlavor
metadata:
name: a100
spec:
nodeLabels:
nvidia.com/gpu.product: NVIDIA-A100-SXM4-80GB # label set by GPU Feature Discovery
---
# ClusterQueue with GPU quota
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
name: gpu-cluster-queue
spec:
namespaceSelector: {}
resourceGroups:
- coveredResources: ["cpu", "memory", "nvidia.com/gpu"]
flavors:
- name: a100
resources:
- name: "nvidia.com/gpu"
nominalQuota: 8
- name: "cpu"
nominalQuota: 64
- name: "memory"
nominalQuota: 256Gi
---
# LocalQueue for team namespace
apiVersion: kueue.x-k8s.io/v1beta2
kind: LocalQueue
metadata:
name: team-ml-queue
namespace: ml-team
spec:
clusterQueue: gpu-cluster-queue
Submit a Job to a Queue¶
apiVersion: batch/v1
kind: Job
metadata:
name: finetune
namespace: ml-team
labels:
kueue.x-k8s.io/queue-name: team-ml-queue # Kueue suspends the Job until quota admits it
spec:
parallelism: 2
completions: 2
suspend: true
template:
spec:
restartPolicy: Never
containers:
- name: train
image: pytorch/pytorch:2.3.0-cuda12.1-cudnn8-runtime
command: ["python", "-c", "import torch; print(torch.cuda.is_available())"]
resources:
limits:
nvidia.com/gpu: 1
Install KAI Scheduler (Alternative)¶
# Pick a version from the releases page (v0.18.0 on 2026-09-23; even minors are LTS)
helm upgrade -i kai-scheduler oci://ghcr.io/kai-scheduler/kai-scheduler/kai-scheduler \
-n kai-scheduler --create-namespace --version <VERSION>
Pods opt in with schedulerName: kai-scheduler and a queue label. See the KAI Scheduler quick start.
GPU Monitoring¶
DCGM Exporter Metrics¶
The DCGM Exporter runs as part of the GPU Operator and exposes Prometheus metrics. The full default metric list is in Reference.
Essential Monitoring Queries (PromQL)¶
# GPU utilization across cluster (coarse)
avg(DCGM_FI_DEV_GPU_UTIL) by (gpu, Hostname)
# Better: fraction of time compute engines were busy, and Tensor Core activity
avg(DCGM_FI_PROF_GR_ENGINE_ACTIVE) by (Hostname)
avg(DCGM_FI_PROF_PIPE_TENSOR_ACTIVE) by (Hostname)
# GPU memory usage percentage
DCGM_FI_DEV_FB_USED / (DCGM_FI_DEV_FB_USED + DCGM_FI_DEV_FB_FREE) * 100
# Idle GPUs (utilization < 5% for 10 minutes)
max_over_time(DCGM_FI_DEV_GPU_UTIL[10m]) < 5
# GPU temperature alerts
DCGM_FI_DEV_GPU_TEMP > 85
# XID errors (non-zero means a driver-reported fault)
DCGM_FI_DEV_XID_ERRORS > 0
# Pods requesting GPUs (kube-state-metrics)
kube_pod_container_resource_limits{resource="nvidia_com_gpu"} > 0
nvidia-smi Quick Reference¶
# Full GPU status
nvidia-smi
# Continuous monitoring (refresh every 1 second)
nvidia-smi -l 1
# Query specific metrics
nvidia-smi --query-gpu=name,temperature.gpu,utilization.gpu,utilization.memory,memory.used,memory.total --format=csv
# Show running GPU processes
nvidia-smi pmon -s um -d 1
# Check MIG status
nvidia-smi mig -lgi -i 0
# Show NVLink status
nvidia-smi nvlink -s
# Show GPU topology
nvidia-smi topo -m
Secure GPU Workloads¶
The threat model is in Explanation. Checklists and port tables are in Reference.
Kubernetes RBAC for GPU Resources¶
Control who can schedule GPU workloads to prevent unauthorized GPU consumption:
# Role restricting GPU pod creation
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: gpu-user
namespace: ml-team
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["create", "get", "list", "delete"]
- apiGroups: ["batch"]
resources: ["jobs"]
verbs: ["create", "get", "list", "delete"]
---
# ResourceQuota limiting GPU allocation per namespace
apiVersion: v1
kind: ResourceQuota
metadata:
name: gpu-quota
namespace: ml-team
spec:
hard:
requests.nvidia.com/gpu: "4"
limits.nvidia.com/gpu: "4"
With DRA, quota applies to device requests instead: ResourceQuota supports <device-class-name>.deviceclass.resource.k8s.io/devices (for example gpu.nvidia.com.deviceclass.resource.k8s.io/devices: "4").
Restrict Access to vLLM¶
vLLM's --api-key gives one static bearer token and nothing more (no per-user identity, quotas or audit). Put a gateway in front for token validation, rate limiting and logging, and restrict pod ingress:
# NetworkPolicy: only allow inference gateway to reach vLLM
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: vllm-access
namespace: inference
spec:
podSelector:
matchLabels:
app: vllm
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
app: inference-gateway
ports:
- port: 8000
Isolate a Ray Cluster¶
# NetworkPolicy: isolate Ray cluster
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: ray-cluster-isolation
namespace: ray
spec:
podSelector:
matchLabels:
ray.io/cluster: gpu-cluster
policyTypes:
- Ingress
- Egress
ingress:
- from:
- podSelector:
matchLabels:
ray.io/cluster: gpu-cluster
- from:
- podSelector:
matchLabels:
app: ray-job-submitter
ports:
- port: 8265
egress:
- to:
- podSelector:
matchLabels:
ray.io/cluster: gpu-cluster
- to: [] # Allow outbound for model downloads (narrow this in production)
Dedicate GPU Nodes to a Team¶
# Taint GPU nodes for a specific team
kubectl taint nodes gpu-node-1 team=ml-production:NoSchedule
kubectl label nodes gpu-node-1 gpu-pool=production
apiVersion: v1
kind: Pod
metadata:
name: production-inference
namespace: ml-production
spec:
tolerations:
- key: "team"
operator: "Equal"
value: "ml-production"
effect: "NoSchedule"
nodeSelector:
gpu-pool: production
containers:
- name: vllm
image: vllm/vllm-openai:v0.30.0
resources:
limits:
nvidia.com/gpu: 1
Mount Model Registry Credentials¶
# Secret for model registry credentials (create it with kubectl, never commit it)
# kubectl create secret generic hf-token -n inference --from-literal=token=<token>
apiVersion: v1
kind: Pod
metadata:
name: vllm
namespace: inference
spec:
containers:
- name: vllm
image: vllm/vllm-openai:v0.30.0
env:
- name: HF_TOKEN
valueFrom:
secretKeyRef:
name: hf-token
key: token
- name: HF_HOME
value: "/models/cache"
volumeMounts:
- name: model-cache
mountPath: /models/cache
volumes:
- name: model-cache
persistentVolumeClaim:
claimName: model-cache-pvc
Detect GPU Cryptomining¶
# Alert on sustained high GPU utilization from pods without an approved workload label
DCGM_FI_DEV_GPU_UTIL > 90
unless on (namespace, pod)
kube_pod_labels{label_workload_type=~"training|inference"}
Pair the alert with prevention: Pod Security Standards (restricted) on GPU namespaces, image allowlisting, admission policies (ValidatingAdmissionPolicy or Kyverno) that only admit GPU requests from approved patterns, and per-namespace GPU quotas.
Audit GPU Resource Events¶
# Kubernetes audit policy for GPU resource events
apiVersion: audit.k8s.io/v1
kind: Policy
rules:
- level: Metadata
resources:
- group: ""
resources: ["pods"]
- group: "resource.k8s.io"
resources: ["resourceclaims"]
verbs: ["create", "delete"]
# Log who creates/deletes GPU pods and DRA claims
- level: RequestResponse
resources:
- group: "batch.volcano.sh"
resources: ["jobs"]
# Full audit for Volcano jobs
Troubleshooting¶
GPU Not Visible to Kubernetes¶
# Check if driver is loaded
lsmod | grep nvidia
# Check device plugin pods
kubectl get pods -n gpu-operator -l app=nvidia-device-plugin-daemonset
# Check device plugin logs
kubectl logs -n gpu-operator -l app=nvidia-device-plugin-daemonset
# Check the operator's own validation
kubectl get pods -n gpu-operator -l app=nvidia-operator-validator
# Verify extended resources on node
kubectl describe node <node> | grep -A5 "Allocatable"
Pod Stuck in Pending (GPU)¶
# Check if GPU resources are available
kubectl describe node <node> | grep nvidia.com/gpu
# Check events on the pending pod
kubectl describe pod <pod-name> | tail -20
# If Kueue is installed: is the workload admitted?
kubectl get workloads -n <namespace>
# Common causes:
# - No GPU nodes available, or all GPUs allocated
# - Resource request exceeds node capacity
# - Taints/tolerations preventing scheduling
# - Queue quota exhausted (Kueue/Volcano/KAI) or gang minimum not reachable
GPU Out of Memory (OOM)¶
# Check GPU memory usage
nvidia-smi
# For vLLM: lower --gpu-memory-utilization (default 0.92 in v0.20+),
# lower --max-model-len or --max-num-seqs, or use --kv-cache-dtype fp8
# For training: reduce batch size or enable gradient checkpointing
# torch.cuda.empty_cache() to free cached memory
Ray Cluster Issues¶
# Check Ray head logs
kubectl logs <ray-head-pod> -c ray-head
# Access Ray dashboard
kubectl port-forward svc/gpu-cluster-head-svc 8265:8265
# Check cluster resources (run inside the head pod; ray status talks to GCS, not the dashboard)
kubectl exec -it <ray-head-pod> -c ray-head -- ray status
# Check autoscaler logs (sidecar added by enableInTreeAutoscaling)
kubectl logs <ray-head-pod> -c autoscaler
Best Practices¶
GPU Resource Management¶
- Right-size GPU requests — Profile workloads before choosing GPU allocation. Use
nvidia-smi pmonandDCGM_FI_PROF_GR_ENGINE_ACTIVEto measure actual utilization. - Use MIG for multi-tenant clusters — Hardware isolation prevents noisy-neighbor issues.
- Enable time-slicing for development — Allow multiple dev workloads to share GPUs. Reserve dedicated GPUs for production inference.
- Set GPU memory limits in vLLM — Keep
--gpu-memory-utilizationaround 0.85-0.92 to leave headroom for CUDA context and spikes. Lower it when other processes share the GPU. - Monitor GPU utilization continuously — A common target is above 70% for production inference. Investigate anything below 50%. These thresholds are rules of thumb, not vendor guidance.
Scheduling¶
- Use Kueue for admission control — Prevent cluster overcommit by queuing jobs that exceed available GPU capacity.
- Use a gang-capable scheduler for distributed training — Volcano, KAI Scheduler, Kueue with
waitForPodsReady, or the native Workload API (beta in 1.37) prevent partial allocation waste. - Enable topology awareness — For multi-GPU training, prefer same-node placement to use NVLink (Kueue TAS, KAI TAS, Volcano network topology).
- Set preemption policies — Allow production inference to preempt batch training during capacity pressure.
Inference Optimization¶
- Use an optimized engine for LLM serving — vLLM (or SGLang, TensorRT-LLM) with PagedAttention and continuous batching. The PagedAttention paper reports 2-4x throughput over FasterTransformer and Orca at the same latency.
- Enable tensor parallelism — For models that exceed single GPU memory, split across GPUs with
--tensor-parallel-size. - Quantize models — FP8, AWQ or GPTQ checkpoints roughly halve weight memory versus FP16/BF16 with small quality loss. Validate on your evals.
- Route by KV cache, not round-robin — At multi-replica scale, prefix-cache-aware routing (Gateway API Inference Extension, llm-d, Dynamo router) cuts TTFT and raises throughput.
- Consider prefill/decode disaggregation for large models — llm-d and Dynamo split prefill and decode pools. Gains depend on workload shape (long prompts, large models, fast interconnect).