How-to Guides¶
Scope
Tasks for running upstream Kubernetes: planning a production cluster, upgrading a kubeadm cluster to v1.37, clearing the v1.37 upgrade blockers, moving off Ingress NGINX, tuning, securing, monitoring, backing up, and day-to-day kubectl recipes. For version dates, skew rules, and removals, see Reference. For how the parts fit together, see Explanation.
Cluster Architecture Patterns¶
Control Plane High Availability¶
| Pattern | etcd Topology | API Server | Min Nodes | Use Case |
|---|---|---|---|---|
| Stacked | Co-located with control plane | 3+ | 3 | Most deployments |
| External | Dedicated etcd cluster | 3+ | 6 (3 etcd + 3 CP) | Enterprise, large scale |
| Single | Single node | 1 | 1 | Dev/test only |
Put a load balancer (or kube-vip / keepalived + HAProxy) in front of the API servers on TCP 6443. Pass its address as --control-plane-endpoint to kubeadm init.
Node Sizing Guidelines¶
Rule of thumb, not an upstream recommendation
Size nodes from your own workload's requests. Remember that upstream tests at most min(110, 10 x cores) pods per node (see Reference).
| Workload Type | vCPUs | Memory | Storage | Network |
|---|---|---|---|---|
| General purpose | 4-8 | 16-32Gi | 100Gi SSD | 10Gbps |
| Memory-intensive (DB) | 8-16 | 64-128Gi | 500Gi NVMe | 10Gbps |
| GPU/ML | 8+ + GPU | 64Gi+ | 1Ti NVMe | 25Gbps |
| Edge/IoT | 2 | 4Gi | 32Gi | 1Gbps |
Node Prerequisites (v1.35+)¶
- cgroup v2 only: since v1.35 the kubelet refuses to start on cgroup v1 unless
failCgroupV1: falseis set. Check withstat -fc %T /sys/fs/cgroup/, which should printcgroup2fs. - Container runtime: containerd 2.3+ or 2.4+ for v1.37 (see Reference), or the matching CRI-O minor.
- Swap: supported on cgroup v2 with
memorySwap.swapBehavior: LimitedSwapin the kubelet config. Otherwise keep swap off.
Upgrade Procedures¶
Cluster Version Upgrade¶
Version Skew Policy
The kubelet may be up to three minor versions older than kube-apiserver and never newer. Controller-manager and scheduler may be at most one minor version older. Always upgrade the control plane first, then the nodes. Do not skip minor versions. See Reference.
The next flowchart shows the kubeadm upgrade order for one minor-version step.
flowchart TD
A["Point pkgs.k8s.io repo at v1.37<br/>on every node"] --> B["First control plane node:<br/>upgrade kubeadm package"]
B --> C["kubeadm upgrade plan"]
C --> D["kubeadm upgrade apply v1.37.x"]
D --> E["Other control plane nodes:<br/>kubeadm upgrade node"]
E --> F["Per node: kubectl drain"]
F --> G["Upgrade kubelet + kubectl packages,<br/>systemctl restart kubelet"]
G --> H["kubectl uncordon"]
H --> I{"More nodes?"}
I -->|yes| F
I -->|no| J["Upgrade CNI, CSI, add-ons<br/>per their own docs"]
The steps below follow the official kubeadm upgrade guide for Debian/Ubuntu with the community pkgs.k8s.io repositories. Each repository serves exactly one minor version, so you must switch it first.
# 0. On every node: switch the apt repo to the target minor (v1.36 -> v1.37)
sudo sed -i 's|/v1.36/|/v1.37/|' /etc/apt/sources.list.d/kubernetes.list
sudo apt-get update
sudo apt-cache madison kubeadm # find the latest 1.37.x-* package
# 1. First control plane node: upgrade kubeadm, then plan and apply
sudo apt-mark unhold kubeadm && \
sudo apt-get install -y kubeadm='1.37.1-*' && \
sudo apt-mark hold kubeadm
sudo kubeadm upgrade plan
sudo kubeadm upgrade apply v1.37.1
# 2. Remaining control plane nodes (after upgrading kubeadm the same way)
sudo kubeadm upgrade node
# 3. Every node, one at a time (control plane nodes too): kubelet + kubectl
kubectl drain <node> --ignore-daemonsets --delete-emptydir-data
sudo apt-mark unhold kubelet kubectl && \
sudo apt-get install -y kubelet='1.37.1-*' kubectl='1.37.1-*' && \
sudo apt-mark hold kubelet kubectl
sudo systemctl daemon-reload && sudo systemctl restart kubelet
kubectl uncordon <node>
On worker nodes, also run sudo kubeadm upgrade node (after upgrading the kubeadm package) before the kubelet step. It refreshes the local kubelet configuration.
Rolling Upgrade Strategy¶
- Take an etcd snapshot (see etcd Backup).
- Upgrade control plane nodes one at a time.
- Upgrade worker nodes in batches (10-20% at a time). With PodDisruptionBudgets,
kubectl drainrespects availability. - Validate workload health between batches.
- On managed or cloud clusters, prefer blue/green node pools: create a new pool at the new version, cordon and drain the old pool, and keep it until the new one is validated.
Prepare for the v1.37 Upgrade¶
These are the "action required" items from CHANGELOG-1.37.md, plus the removals from earlier releases that still catch clusters out.
- SELinux hosts:
SELinuxMountis GA and on by default. Volumes are mounted with-o context=...instead of relabelled recursively. This can break pods that share a volume while using different SELinux labels. Find affected workloads on v1.36 first (Kubernetes blog, 2026-04-22, "breaking changes in SELinux volume labeling"). Then fix them or opt out per pod withseLinuxChangePolicy: Recursive. - Workload-aware scheduling alpha users: delete every
scheduling.k8s.io/v1alpha2Workload/PodGroupobject before you upgrade.v1alpha2is dropped. - cgroup v1 nodes: migrate to cgroup v2. The
failCgroupV1: falseopt-out still works but is deprecated. - kube-proxy mode: set
mode:explicitly inKubeProxyConfiguration. Plan anipvs->nftablesmigration (needs a recent kernel. iptables remains the fallback). - containerd: move to 2.3+/2.4+. containerd 1.7 is outside the recommended matrix from 1.36 on.
- kubelet flags: remove deprecated cAdvisor flags (the kubelet fails to start if they are set) and
--pod-infra-container-image(removed in 1.35). - kubelet
eventRecordQPS: 0now means unlimited. Set an explicit value (for example50) if you relied on the old behaviour. - RBAC: restrict
nodes/logs. The 1.37 kubelet logs its effective configuration at startup. - PodCertificateRequest clients: stop setting
PKIXPublicKey/ProofOfPossession. They are removed fromv1.
Migrate from Ingress NGINX to Gateway API¶
Ingress NGINX (kubernetes/ingress-nginx) was retired in March 2026 and gets no further security fixes. The steps below move you to a Gateway API implementation (for example Envoy Gateway, Istio, Cilium, NGINX Gateway Fabric, or a cloud provider's controller).
# 1. Install the Gateway API CRDs (standard channel) - pin the version you tested
kubectl apply --server-side -f \
https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.6.2/standard-install.yaml
# 2. Install a controller, e.g. Envoy Gateway (see its docs for the current chart version)
helm install eg oci://docker.io/envoyproxy/gateway-helm --version v1.9.1 -n envoy-gateway-system --create-namespace
# 3. Convert existing Ingress objects to Gateway API resources
# (ingress2gateway reads ingress-nginx annotations and prints Gateway/HTTPRoute YAML)
go install github.com/kubernetes-sigs/ingress2gateway@v1.0.0
ingress2gateway print --providers=ingress-nginx --emitter=envoy-gateway -A > gateway-resources.yaml
Then review the output by hand. Annotations with no Gateway API equivalent (snippets, custom Lua) need a redesign. Run both paths in parallel behind DNS, then move traffic over. For controller choices, see Envoy Gateway and Service Mesh Comparison.
Performance Tuning¶
API Server¶
# kube-apiserver flags (static pod manifest: /etc/kubernetes/manifests/kube-apiserver.yaml)
--max-requests-inflight=400 # default 400 (non-mutating)
--max-mutating-requests-inflight=200 # default 200
# With API Priority and Fairness (GA since v1.29) the two limits above form the
# total concurrency budget that APF divides between PriorityLevelConfigurations.
--watch-cache-sizes=events#0 # format resource[.group]#size; only 0 (disable) is meaningful,
# watch cache sizing is otherwise automatic
Tune fairness with FlowSchema and PriorityLevelConfiguration objects (flowcontrol.apiserver.k8s.io/v1) rather than by raising the in-flight limits.
etcd¶
| Parameter | Small (< 100 nodes) | Large (100+ nodes) |
|---|---|---|
--quota-backend-bytes |
2Gi (default) | 8Gi (etcd's suggested maximum) |
--snapshot-count |
10000 | 50000 |
--auto-compaction-retention |
1h | 5m |
| Storage | SSD (min) | NVMe (required) |
kube-apiserver already requests compaction every 5 minutes (--etcd-compaction-interval, default 5m). Defragment each member during maintenance windows, one member at a time.
etcd Performance
etcd is the most common bottleneck in large clusters. Use dedicated SSD/NVMe storage with a p99 WAL fsync under 10 ms. Run etcdctl check perf against a test cluster (it generates load).
Kubelet¶
# /var/lib/kubelet/config.yaml
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
maxPods: 110 # default. Upstream tests up to min(110, 10 x cores)
containerLogMaxSize: "50Mi"
containerLogMaxFiles: 5
imageGCHighThresholdPercent: 85
imageGCLowThresholdPercent: 80
evictionHard:
memory.available: "500Mi"
nodefs.available: "10%"
imagefs.available: "15%"
On kubeadm clusters, edit the kube-system/kubelet-config ConfigMap and run kubeadm upgrade node phase kubelet-config on each node, so the change survives upgrades.
Resize Pods In Place¶
In-place Pod resize is GA since v1.35. Change CPU and memory on a running pod through the resize subresource without recreating it:
kubectl patch pod myapp-xxx --subresource resize --patch \
'{"spec":{"containers":[{"name":"app","resources":{"requests":{"cpu":"800m"},"limits":{"cpu":"800m"}}}]}}'
kubectl get pod myapp-xxx -o jsonpath='{.status.conditions}' # PodResizePending / PodResizeInProgress
Secure the Cluster¶
Enforce Pod Security Standards¶
# Apply the restricted profile to a namespace (Pod Security Admission, GA since v1.25)
kubectl label namespace production \
pod-security.kubernetes.io/enforce=restricted \
pod-security.kubernetes.io/enforce-version=latest \
pod-security.kubernetes.io/audit=restricted \
pod-security.kubernetes.io/warn=restricted
Encryption at Rest¶
Enable encryption by pointing kube-apiserver --encryption-provider-config at an EncryptionConfiguration. The first provider in the list encrypts new writes. The others are only used to decrypt.
apiVersion: apiserver.config.k8s.io/v1
kind: EncryptionConfiguration
resources:
- resources:
- secrets
providers:
- kms: # KMS v2 (GA since v1.29) - recommended
apiVersion: v2
name: my-kms-plugin
endpoint: unix:///var/run/kms-plugin/socket.sock
- secretbox: # strong local alternative if no KMS is available
keys:
- name: key1
secret: <BASE64_ENCODED_32_BYTE_KEY>
- identity: {} # allows reading data written before encryption was enabled
Provider choice
aescbc is not recommended (CBC is vulnerable to padding-oracle attacks). aesgcm is not recommended unless you rotate keys automatically. KMS v1 has been deprecated since v1.28. After you enable encryption, rewrite existing Secrets: kubectl get secrets -A -o json | kubectl replace -f -. See Encrypting Secret Data at Rest.
Enable Audit Logging¶
# /etc/kubernetes/audit-policy.yaml, passed with --audit-policy-file and --audit-log-path
apiVersion: audit.k8s.io/v1
kind: Policy
rules:
- level: Metadata # never log Secret bodies
resources:
- group: ""
resources: ["secrets", "configmaps"]
- level: RequestResponse
resources:
- group: "rbac.authorization.k8s.io"
- level: None
users: ["system:kube-proxy"]
- level: Metadata
Do not log Secret payloads
A RequestResponse rule on secrets writes Secret values into the audit log in plain text. Log Secrets at Metadata level.
Scan Against the CIS Benchmark¶
# kube-bench runs the CIS Kubernetes Benchmark checks as a Job
kubectl apply -f https://raw.githubusercontent.com/aquasecurity/kube-bench/main/job.yaml
kubectl logs job/kube-bench
Common Issues & Troubleshooting¶
| Symptom | Diagnosis | Resolution |
|---|---|---|
Node NotReady |
kubectl describe node, journalctl -u kubelet |
Check kubelet logs, disk pressure, memory, CNI |
| kubelet fails to start after upgrade to 1.35+ | journalctl -u kubelet mentions cgroup v1 |
Move the node to cgroup v2, or set failCgroupV1: false as a stop-gap |
| kubelet fails to start on 1.37 | Unknown cAdvisor flag in logs | Remove the deprecated cAdvisor flags |
Pod stuck Pending |
kubectl describe pod |
Check resource requests, node affinity, taints, PVC binding, DRA claims |
| DNS resolution failing | kubectl exec -it <pod> -- nslookup kubernetes.default |
Check CoreDNS pods and logs, resolv.conf, NetworkPolicy to kube-dns port 53 |
| CrashLoopBackOff | kubectl logs <pod> --previous |
Fix application error, check probes |
| ImagePullBackOff | kubectl describe pod events |
Verify image and tag exist, check pull secrets and registry auth |
| etcd leader changes | etcdctl endpoint status -w table |
Check disk fsync latency, network partitions |
| Pod resize stuck | kubectl get pod -o yaml conditions PodResizePending |
Node lacks capacity (Deferred) or the request is impossible (Infeasible) |
Monitoring Stack¶
Essential Metrics¶
# Cluster CPU request utilisation (kube-state-metrics)
sum(kube_pod_container_resource_requests{resource="cpu"})
/ sum(kube_node_status_allocatable{resource="cpu"})
# Node memory pressure (node_exporter)
node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes < 0.1
# API server p99 latency by verb (excluding long-running requests)
histogram_quantile(0.99,
sum by (le, verb) (rate(apiserver_request_duration_seconds_bucket{verb!~"WATCH|CONNECT"}[5m])))
# etcd WAL fsync p99 (should be < 10ms)
histogram_quantile(0.99, sum by (le, instance) (rate(etcd_disk_wal_fsync_duration_seconds_bucket[5m])))
# Pods restarting in the last hour
increase(kube_pod_container_status_restarts_total[1h]) > 0
For a packaged stack, install kube-prometheus-stack (see the Helm recipe below) or use VictoriaMetrics.
Backup & Disaster Recovery¶
etcd Backup¶
# Snapshot etcd (etcdctl v3 API is the default since etcd 3.4)
etcdctl snapshot save /backup/etcd-$(date +%Y%m%d).db \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/healthcheck-client.crt \
--key=/etc/kubernetes/pki/etcd/healthcheck-client.key
# Verify snapshot (etcd 3.5+: use etcdutl for offline file operations)
etcdutl snapshot status /backup/etcd-$(date +%Y%m%d).db --write-out=table
Suggested practice: snapshot every 6-12 hours and before every upgrade. Copy snapshots off-cluster to object storage. Keep several generations. Rehearse etcdutl snapshot restore on a scratch cluster.
Velero for Workload Backup¶
velero backup create full-backup --include-namespaces '*'
velero restore create --from-backup full-backup
Commands & Recipes¶
Cluster Operations¶
# Cluster info
kubectl cluster-info
kubectl get nodes -o wide
kubectl top nodes # needs metrics-server (metrics.k8s.io/v1 since v1.37)
# Drain node for maintenance
kubectl drain node-1 --ignore-daemonsets --delete-emptydir-data
kubectl uncordon node-1
# Check control plane health (componentstatuses is deprecated since v1.19)
kubectl get --raw '/readyz?verbose'
kubectl get --raw '/livez?verbose'
Workload Management¶
# Deploy and scale
kubectl apply -f deployment.yaml
kubectl scale deployment myapp --replicas=5
kubectl rollout status deployment myapp
# Rolling update
kubectl set image deployment/myapp app=myapp:v2.0
kubectl rollout undo deployment/myapp # rollback
# Restart pods (rolling)
kubectl rollout restart deployment/myapp
# Port forward for debugging
kubectl port-forward svc/myapp 8080:80
# Run one-off debug pod
kubectl run debug --rm -it --image=nicolaka/netshoot -- /bin/bash
# KYAML output (stable in v1.37): stricter, less ambiguous YAML
kubectl get deployment myapp -o kyaml
Debugging¶
# Pod debugging
kubectl describe pod myapp-xxx # events, conditions
kubectl logs myapp-xxx -c app --previous # previous crash logs
kubectl logs -l app=myapp --all-containers -f # follow all pods
# Exec into running pod
kubectl exec -it myapp-xxx -- /bin/sh
# Ephemeral debug container in a running pod (distroless images)
kubectl debug -it myapp-xxx --image=busybox:1.36 --target=app
# Debug node
kubectl debug node/node-1 -it --image=ubuntu
# Check events (sorted by time)
kubectl get events --sort-by='.lastTimestamp' -A
# Resource usage
kubectl top pods --sort-by=cpu -A
kubectl top pods --sort-by=memory
Networking¶
# View services and their EndpointSlices (v1 Endpoints is deprecated since v1.33)
kubectl get svc -o wide
kubectl get endpointslices -l kubernetes.io/service-name=myapp
# DNS debugging
kubectl run dns-test --rm -it --image=busybox:1.36 -- nslookup myapp.default.svc.cluster.local
# View network policies and Gateway API routes
kubectl get networkpolicies -A
kubectl get gateways,httproutes -A
RBAC¶
# Check permissions
kubectl auth can-i create pods --namespace=production
kubectl auth can-i '*' '*' --all-namespaces # am I cluster-admin?
kubectl auth whoami # show the identity the API server sees
# Create service account with role
kubectl create serviceaccount deployer
kubectl create rolebinding deployer-binding \
--clusterrole=edit \
--serviceaccount=default:deployer
# Short-lived token for a service account (no Secret-based token needed)
kubectl create token deployer --duration=1h
# Audit broad bindings
kubectl get rolebindings,clusterrolebindings -A -o wide
Helm¶
# Install a chart (example: kube-prometheus-stack)
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
helm install monitoring prometheus-community/kube-prometheus-stack -n monitoring --create-namespace
# Upgrade with values
helm upgrade monitoring prometheus-community/kube-prometheus-stack -n monitoring -f values.yaml
# Rollback to revision 1
helm rollback monitoring 1 -n monitoring
# Template (dry-run render)
helm template monitoring prometheus-community/kube-prometheus-stack -f values.yaml > rendered.yaml
Bitnami charts
Earlier versions of this page used bitnami/postgresql. Bitnami changed how it distributes free images and charts in 2025. Check its current terms before you depend on those charts.
Manifest Patterns¶
# Deployment with resource limits, probes, topology spread and a native sidecar
apiVersion: apps/v1
kind: Deployment
metadata:
name: myapp
spec:
replicas: 3
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: myapp
initContainers:
- name: log-tailer # native sidecar (GA v1.33): restartPolicy Always
image: busybox:1.36
command: ["sh", "-c", "tail -F /var/log/app/app.log"]
restartPolicy: Always
containers:
- name: app
image: myapp:1.4.2 # pin a version or digest, never :latest
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 512Mi
livenessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 5
periodSeconds: 10
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 3
periodSeconds: 5
Sources¶
- kubectl Quick Reference
- kubectl Reference
- Upgrading kubeadm clusters
- Changing the Kubernetes package repository
- CHANGELOG-1.37 (Urgent Upgrade Notes)
- Resize CPU and Memory Resources assigned to Containers
- Encrypting Secret Data at Rest
- Auditing
- Operating etcd clusters for Kubernetes
- ingress2gateway
- Ingress NGINX Retirement