Explanation¶
What this page explains
Kubernetes has a declarative, state-driven architecture. A cluster has a control plane (the API server, etcd, the scheduler, and the controller managers) and worker nodes (kubelet, kube-proxy, and a CRI runtime). All state lives in etcd, and every component talks to the others only through the API server. This page covers why it works this way: the components, the reconciliation loop, pod lifecycle, networking, storage, scheduling, the 2025-2026 design shifts (DRA, in-place resize, cgroup v2, nftables, Gateway API), the release model, and the security model. Numbers and dates are in Reference. Tasks are in How-to Guides.
See also: Kubernetes hub, Reference, How-to Guides
Cluster Architecture Overview¶
The next diagram shows the control plane, two worker nodes, and the cluster DNS add-on. Only kube-apiserver talks to etcd.
graph TD
subgraph ControlPlane["Control Plane"]
API["kube-apiserver<br/>:6443"]
ETCD["etcd<br/>:2379-2380"]
SCHED["kube-scheduler<br/>:10259"]
CM["kube-controller-manager<br/>:10257"]
CCM["cloud-controller-manager<br/>(cloud only)"]
end
subgraph Addons["Cluster add-ons (run as pods)"]
DNS["CoreDNS<br/>(Service kube-dns)"]
CNIA["CNI agent DaemonSet<br/>(Cilium / Calico / Flannel)"]
end
subgraph Node1["Worker Node 1"]
KUBELET1["kubelet<br/>:10250"]
PROXY1["kube-proxy<br/>:10256"]
CRI1["CRI runtime<br/>(containerd / CRI-O)"]
POD1A["Pod A"]
POD1B["Pod B"]
end
subgraph Node2["Worker Node 2"]
KUBELET2["kubelet<br/>:10250"]
PROXY2["kube-proxy<br/>:10256"]
CRI2["CRI runtime<br/>(containerd / CRI-O)"]
POD2A["Pod C"]
POD2B["Pod D"]
end
API <--> ETCD
SCHED -->|"watch + bind"| API
CM -->|"watch + reconcile"| API
CCM -->|"nodes, routes, LBs"| API
DNS -->|"watch Services"| API
KUBELET1 -->|"watch Pods, report status"| API
PROXY1 -->|"watch EndpointSlices"| API
KUBELET2 --> API
PROXY2 --> API
KUBELET1 -->|"CRI gRPC"| CRI1
CRI1 --> POD1A
CRI1 --> POD1B
KUBELET2 -->|"CRI gRPC"| CRI2
CRI2 --> POD2A
CRI2 --> POD2B
Control Plane Components¶
kube-apiserver¶
The API server is the front end of the Kubernetes control plane. It exposes the Kubernetes HTTP API on port 6443 and is the only component that talks to etcd directly.
- Validates and stores all API objects (pods, services, deployments, custom resources)
- Serves the REST API used by kubectl, controllers, and external tools (JSON, Protobuf, and CBOR as an opt-in)
- Runs authentication, authorization (RBAC, Node, webhook), and admission (built-in plugins, CEL
ValidatingAdmissionPolicy/MutatingAdmissionPolicy, webhooks) - Scales horizontally: run several replicas behind a load balancer
- Keeps a watch cache per resource, so the other components watch the API server, not etcd. Recent releases lower etcd load further: consistent reads from cache (GA 1.34), watch-cache initialization via streamed etcd range reads (
EtcdRangeStream, beta 1.37), and gzip-compressed WatchList responses (beta 1.37) - API Priority and Fairness (APF) shares concurrency across tenants so one noisy client cannot starve the rest
etcd¶
etcd is a distributed, strongly consistent key-value store that holds all cluster state:
- Stores every object definition, the cluster configuration, and dynamic state
- Uses the Raft consensus algorithm for HA (a quorum of 3 or 5 members)
- Listens on TCP 2379 (clients) and 2380 (peers)
- Only the API server talks to etcd directly
- Kubernetes v1.37 ships etcd 3.7.0 as its default version
- Backups are critical. Take regular
etcdctl snapshot savesnapshots (see How-to Guides)
etcd Quorum
An etcd cluster needs a majority (quorum) to accept writes. A 3-member cluster survives 1 failure, and a 5-member cluster survives 2. Never run production etcd with fewer than 3 members. Even member counts add no fault tolerance.
kube-scheduler¶
The scheduler watches for new pods that have no node assigned and picks a node for each one. It is built on the scheduling framework, which is a set of plugin extension points:
- PreFilter / Filter: remove nodes that cannot run the pod (resources, taints, node affinity, volume topology, DRA device availability)
- PostFilter: preemption. If no node fits, lower-priority pods may be evicted
- PreScore / Score: rank the remaining nodes with plugins such as
NodeResourcesFit,NodeResourcesBalancedAllocation,InterPodAffinity,PodTopologySpread,TaintToleration, andImageLocality - Reserve / Permit / PreBind / Bind: reserve resources (including DRA allocations) and write
spec.nodeNamethrough thepods/bindingsubresource - QueueingHints (GA 1.34) requeue unschedulable pods only when a relevant cluster event happens
- Workload-aware scheduling (APIs beta in 1.37): the
WorkloadandPodGroupAPIs let the scheduler place a gang of pods all at once (for AI/ML training and MPI), with workload-aware preemption - Serves health and metrics on port 10259
kube-controller-manager¶
Runs the core reconciliation controllers in a single binary:
- Node lifecycle controller: watches node heartbeats (Lease objects) and taints and evicts pods from unreachable nodes
- ReplicaSet controller: keeps the desired number of pod replicas
- Deployment controller: manages rolling updates and rollbacks through ReplicaSets
- StatefulSet controller: ordered, identity-preserving pods. 1.37 adds an alpha
Recreateupdate strategy - Job / CronJob controllers: batch completion, success policies, per-index backoff
- EndpointSlice controller: publishes Service backends (the v1 Endpoints API is deprecated since 1.33)
- HorizontalPodAutoscaler controller: scale to and from zero is on by default since 1.37
- Service account and token controllers, Namespace controller (ordered namespace deletion is GA since 1.34), garbage collector
Each controller is a reconciliation loop. It watches the API server through shared informers and drives the current state toward the desired state.
cloud-controller-manager¶
This component holds the cloud-specific control loops: node initialization and addresses, routes, and type: LoadBalancer Services. They were split out of core Kubernetes so providers can release on their own schedule. In-tree cloud providers were removed in v1.31.
CoreDNS (cluster add-on)¶
CoreDNS provides cluster-wide DNS resolution for Services and Pods. It is not a control plane binary. It is an add-on that runs as a Deployment in kube-system behind a Service still named kube-dns for compatibility:
- Resolves Service names to ClusterIPs (for example,
my-service.my-namespace.svc.cluster.local) - Returns pod IPs directly for headless Services
- Is configured through a Corefile (custom zones, forwarding, rewrites)
- Has been the default since v1.13. The legacy kube-dns add-on is deprecated as of v1.37
Node Components¶
kubelet¶
The kubelet runs on every node and manages the pod lifecycle:
- Watches the API server for pods assigned to its node
- Pulls images and starts containers through the CRI runtime
- Reports node and pod status back to the API server, plus a node Lease heartbeat
- Runs liveness, readiness, and startup probes. 1.37 adds HTTP/2 cleartext (
protocolfield) probes - Mounts volumes (via CSI) and has the runtime set up pod networking (via CNI)
- Applies in-place resizes of CPU and memory to running containers (GA 1.35)
- Listens on port 10250 for the API server (exec, logs, port-forward, metrics)
- Requires cgroup v2 by default since v1.35. It gets its cgroup driver from the runtime over CRI (
KubeletCgroupDriverFromCRI, GA 1.34)
kube-proxy¶
kube-proxy programs each node's dataplane so that Service virtual IPs reach backend pods:
- Watches Services and EndpointSlices
- Modes on Linux: iptables (still the default), nftables (GA 1.33, the planned future default), and ipvs (deprecated since 1.35, with removal planned). The old
userspacemode was removed in 1.26. Windows useskernelspace - Some CNIs (for example Cilium's kube-proxy replacement) do the job of kube-proxy with eBPF instead
- Serves a health check on port 10256
CRI Runtime (containerd / CRI-O)¶
The Container Runtime Interface (CRI) decouples the kubelet from any specific runtime:
- containerd: the most common runtime and a CNCF graduated project. Kubernetes 1.36 and 1.37 are tested with containerd 2.x (2.2+/2.3+/2.4+). containerd 1.7 is outside the recommended matrix from 1.36 on
- CRI-O: a lightweight runtime built only for Kubernetes. Its minor versions track Kubernetes minors
- Both implement the CRI gRPC API that the kubelet calls
- Image pulls, container creation, and execution are all delegated to the runtime
- Dockershim was removed in Kubernetes v1.24. Images built with Docker still run, because they are OCI images
Container Network Interface (CNI)¶
CNI plugins configure pod networking:
- Each pod gets its own IP address (IPv4, IPv6, or dual-stack)
- Plugins include Cilium, Calico, Flannel, Antrea, the AWS VPC CNI, and GKE Dataplane V2
- The plugin provides pod-to-pod connectivity within and across nodes
- NetworkPolicy enforcement depends on the plugin (Flannel alone does not enforce it)
Container Storage Interface (CSI)¶
CSI is a standard interface for exposing storage to containers:
- Out-of-tree CSI drivers have replaced the in-tree volume plugins. CSI migration is GA for the major cloud plugins, and in-tree code has been removed step by step (for example, Portworx operations are fully redirected to CSI as of v1.36)
- Third-party drivers include AWS EBS CSI, GCE PD CSI, Azure Disk CSI, Ceph CSI, and Longhorn
- Drivers handle provisioning, attachment, mounting, snapshots, cloning, expansion, and
VolumeAttributesClasschanges (GA 1.34)
Container Runtime Interface (CRI)¶
The CRI is the gRPC API between the kubelet and the container runtime:
- Defines two services:
RuntimeService(pod sandbox and container lifecycle) andImageService(image management) - Lets runtimes plug in without changes to kubelet source code
- Supports both container runtimes and sandboxed runtimes (Kata Containers, gVisor) through
RuntimeClass - Keeps growing: 1.37 adds
CheckpointPod/RestorePodRPCs for pod-level checkpoint and restore
Request Flow: Pod Creation¶
The next sequence diagram follows a bare Pod from kubectl apply to a running container.
sequenceDiagram
actor User
participant API as kube-apiserver
participant etcd as etcd
participant Sched as kube-scheduler
participant Kubelet as kubelet
participant CRI as CRI runtime
User->>API: kubectl apply -f pod.yaml
API->>API: authenticate + authorize (RBAC)
API->>API: run admission (mutating, validating, CEL policies, webhooks)
API->>etcd: persist pod spec (nodeName empty)
etcd-->>API: confirmed
API-->>User: 201 Created
Sched->>API: watch for unscheduled pods
API-->>Sched: pod with empty nodeName
Sched->>Sched: filter + score nodes
Sched->>API: POST pods/binding (set nodeName)
API->>etcd: persist binding
Kubelet->>API: watch pods for its node
API-->>Kubelet: pod assigned to this node
Kubelet->>CRI: RunPodSandbox, PullImage, CreateContainer, StartContainer
CRI-->>Kubelet: container started
Kubelet->>API: update pod status (Running)
API->>etcd: persist status
How It Works¶
Desired-state reconciliation, pod lifecycle, networking, storage, and scheduling internals.
Desired-State Reconciliation Loop¶
The next sequence diagram shows how a Deployment becomes pods through a chain of independent controllers. Each controller only reacts to watched changes.
sequenceDiagram
participant User as User / CI
participant API as kube-apiserver
participant ETCD as etcd
participant Ctrl as kube-controller-manager
participant Sched as kube-scheduler
participant KL as kubelet (Node)
participant CRI as containerd
User->>API: kubectl apply -f deployment.yaml
API->>ETCD: Store desired state
Ctrl->>API: Watch: new Deployment
Ctrl->>API: Deployment controller creates ReplicaSet
Ctrl->>API: ReplicaSet controller creates Pods
Sched->>API: Watch: unscheduled Pods
Sched->>Sched: Filter and score nodes (resources, affinity, taints)
Sched->>API: Bind Pod to Node
KL->>API: Watch: Pod assigned to my node
KL->>CRI: Create pod sandbox
CRI->>CRI: Pull image, start containers
KL->>API: Update Pod status: Running
loop Reconciliation
Ctrl->>API: Watch: actual vs desired state
Ctrl->>Ctrl: If replicas below desired, create Pods
Ctrl->>Ctrl: If replicas above desired, delete surplus
end
Pod Lifecycle¶
The next state diagram shows the pod phases. Init containers, including native sidecars, run while the pod is Pending. Readiness is a condition inside Running, not a phase.
stateDiagram-v2
[*] --> Pending: Pod accepted
state Pending {
[*] --> Unscheduled
Unscheduled --> Scheduled: bound to node
Scheduled --> InitContainers: image pull, sandbox
InitContainers --> Starting: init done, sidecars running
}
Pending --> Running: all containers started
Running --> Succeeded: all containers exit 0 (restartPolicy Never or OnFailure)
Running --> Failed: container exits non-zero, not restarted
Running --> Unknown: node unreachable
Unknown --> Running: node recovers
Unknown --> Failed: node lost, pod evicted
Succeeded --> [*]
Failed --> [*]
state Running {
[*] --> NotReady
NotReady --> Ready: readiness probe passes
Ready --> NotReady: readiness probe fails
}
Native sidecars (GA 1.33) are init containers with restartPolicy: Always. They start before the app containers, keep running for the pod's lifetime, and stop after the app containers. This fixes the old problems of Jobs never finishing and of shutdown ordering. In-place resize (GA 1.35) changes a running container's CPU and memory through the resize subresource. The kubelet reports progress with the PodResizePending and PodResizeInProgress conditions, and the pod keeps running.
Networking Model¶
The 4 Networking Rules¶
- Pod-to-Pod: every Pod gets its own IP. All Pods can talk to each other without NAT.
- Pod-to-Service: Services provide stable virtual IPs (ClusterIP), implemented by kube-proxy (iptables, nftables, ipvs) or an eBPF replacement.
- External-to-Service: LoadBalancer, NodePort, or Ingress / Gateway API expose Services.
- Pod-to-External: Pods reach external networks through SNAT.
The next diagram shows those paths for a Service backed by pods on two nodes.
flowchart TB
subgraph Cluster["Kubernetes Cluster"]
subgraph Node1["Node 1"]
P1["Pod A<br/>10.244.1.2"]
P2["Pod B<br/>10.244.1.3"]
KP1["kube-proxy<br/>(iptables / nftables)"]
end
subgraph Node2["Node 2"]
P3["Pod C<br/>10.244.2.2"]
P4["Pod D<br/>10.244.2.3"]
KP2["kube-proxy"]
end
SVC["Service my-svc<br/>ClusterIP 10.96.0.10<br/>backends: Pod A, Pod C"]
GW["Gateway + HTTPRoute<br/>(Gateway API controller)"]
CNI["CNI plugin<br/>(Cilium / Calico / Flannel)<br/>pod-to-pod routing"]
end
External["External traffic"] -->|"LoadBalancer / NodePort"| GW
GW --> SVC
KP1 -.->|"programs rules"| SVC
KP2 -.->|"programs rules"| SVC
SVC -->|"DNAT"| P1
SVC -->|"DNAT"| P3
P1 <-->|"CNI overlay or native routing"| P3
P2 <-->|"CNI"| P4
CNI -.-> P1
CNI -.-> P3
Ingress vs Gateway API¶
Ingress (networking.k8s.io/v1) is frozen: it still works, but new work happens in Gateway API. Gateway API is a set of role-oriented CRDs: the infrastructure provider owns GatewayClass, the cluster operator owns Gateway, and application developers own HTTPRoute, GRPCRoute, TLSRoute, TCPRoute, and UDPRoute. It supports traffic splitting, header matching, and cross-namespace delegation with ReferenceGrant. The widely used Ingress NGINX controller was retired in March 2026. That made Gateway API the default path for new clusters. See How-to Guides.
Storage Architecture¶
The next diagram shows dynamic provisioning from a PVC down to a real disk.
flowchart LR
Pod["Pod"] --> PVC["PersistentVolumeClaim<br/>(request: 10Gi)"]
PVC --> PV["PersistentVolume<br/>(10Gi, RWO)"]
PVC -.->|"storageClassName"| SC["StorageClass<br/>(provisioner: ebs.csi.aws.com)"]
SC --> CSI["CSI driver<br/>(controller + node plugins)"]
CSI -->|"CreateVolume"| PV
CSI --> Disk["Cloud disk or<br/>storage backend"]
Scheduling Algorithm¶
| Phase | Operation |
|---|---|
| Filtering | Remove nodes that do not meet the pod's requirements: resources (NodeResourcesFit), taints, node affinity, volume topology, DRA devices |
| Scoring | Rank the remaining nodes. The default plugins include NodeResourcesFit (LeastAllocated strategy), NodeResourcesBalancedAllocation, NodeAffinity, InterPodAffinity, PodTopologySpread, TaintToleration, and ImageLocality |
| Binding | Assign the pod to the highest-scoring node through the binding subresource |
| Preemption | PostFilter: if no node fits, evict lower-priority pods. 1.37 adds alpha preemption for deferred in-place resizes |
Dynamic Resource Allocation¶
DRA (core GA in v1.34, API resource.k8s.io/v1) replaces the device-plugin model, which could only count opaque integers such as nvidia.com/gpu: 1, with claims for devices that have attributes. Drivers publish ResourceSlice objects that describe devices and their attributes. Admins define DeviceClass objects. Workloads request devices through a ResourceClaim or a ResourceClaimTemplate with CEL selectors. The scheduler allocates matching devices itself, so it knows about devices before it binds the pod. Releases since GA have added prioritized alternatives and admin access (GA 1.36), and device taints/tolerations and extended-resource mapping (GA 1.37). That second feature lets existing nvidia.com/gpu-style requests be served by DRA drivers.
The next sequence diagram shows the DRA allocation flow for a pod that references a claim template.
sequenceDiagram
participant Drv as DRA driver (kubelet plugin)
participant API as kube-apiserver
participant RCC as resourceclaim controller
participant Sched as kube-scheduler
participant KL as kubelet
Drv->>API: publish ResourceSlice (devices + attributes)
Note over API: Pod references ResourceClaimTemplate
RCC->>API: create ResourceClaim for the Pod
Sched->>API: watch Pod, ResourceClaim, ResourceSlices, DeviceClasses
Sched->>Sched: filter nodes by CEL selectors and free devices
Sched->>API: write allocation to ResourceClaim status
Sched->>API: bind Pod to node
KL->>Drv: NodePrepareResources (gRPC)
Drv-->>KL: CDI device IDs
KL->>KL: start containers with CDI devices injected
KL->>Drv: NodeUnprepareResources when the Pod ends
Design Shifts in 2025-2026¶
| Shift | What changed | Why |
|---|---|---|
| Devices as first-class API | DRA core GA (1.34). Feature set filled in through 1.37 | GPUs, NICs, and accelerators need attributes, sharing, and topology that device plugins cannot express |
| Mutable pod resources | In-place resize GA (1.35), including init containers (GA 1.37) and pod-level resources (beta) | Vertical scaling without restarts, for JVMs, AI inference, and VPA |
| Lifecycle-correct sidecars | Native sidecars GA (1.33) | Service mesh and log-agent containers need defined start and stop ordering |
| Linux baseline raised | cgroup v1 refused by default (1.35). User namespaces GA (1.36). Rootless kubelet beta (1.37) | cgroup v2 is what PSI metrics, MemoryQoS, and swap need. User namespaces shrink the container-escape blast radius |
| Dataplane modernisation | nftables GA (1.33). ipvs deprecated (1.35). Default-mode switch announced (1.37) | iptables rule-update cost grows with Service count. ipvs was under-maintained |
| North-south traffic | Gateway API replaces Ingress. Ingress NGINX retired (March 2026) | Ingress was under-specified (annotation sprawl), and the controller lacked maintainers |
| In-tree policy | CEL ValidatingAdmissionPolicy (GA 1.30) and MutatingAdmissionPolicy (GA 1.36) |
Webhooks add latency and failure modes. CEL runs inside the API server |
| AI/ML batch | Workload-aware / gang scheduling APIs (beta 1.37). HPA scale-to-zero on by default (1.37) | Distributed training needs all-or-nothing placement. Idle GPU inference should scale to zero |
Release Cadence and Support Model¶
Kubernetes releases about three minor versions a year, roughly every 15-17 weeks (1.34 in August 2025, 1.35 in December 2025, 1.36 in April 2026, 1.37 in August 2026). The project supports the latest three minors. Each gets about 12 months of patches plus a 2-month maintenance-mode tail, so any version lives about 14 months. Two KEPs set this model: KEP-1498 gave each release a one-year support period (from v1.19), and KEP-2572 moved Kubernetes from four releases a year to three (from 2021). The skew policy lets kubelets lag the API server by three minors. That means node fleets can be upgraded less often than control planes. Control planes still have to step through every minor, one at a time. Managed providers add paid extended support (EKS, GKE Extended channel, AKS LTS) for organisations that cannot keep up with that cadence. Current dates are in Reference.
Scalability and Performance¶
Upstream does not publish "maximums". It publishes a tested envelope: 5,000 nodes, 150,000 pods, and at most 110 pods per node. Inside that envelope, the SIG Scalability SLOs hold, for example p99 mutating API latency of 1 s or less, and p99 stateless pod startup of 5 s or less. Going beyond the envelope is possible but becomes your own engineering problem. The limiting factors are usually etcd write latency and database size, watch fan-out from the API server, and scheduler throughput with complex affinity. Hyperscalers have gone far past the envelope by replacing parts of the stack: GKE supports 65,000-node clusters, and in 2025 it ran an experimental 130,000-node cluster on a Spanner-based store. The thresholds, SLOs, and indicative figures are in Reference.
Security Model¶
Kubernetes security follows the 4C's model: Cloud, Cluster, Container, and Code. This section covers the cluster and container layers: identity, access control, pod isolation, secrets, and supply chain. The setup tasks are in How-to Guides, and the hardening checklist is in Reference.
Authentication and Identity¶
Users and Service Accounts¶
Kubernetes has two kinds of identity:
- Users: managed outside the cluster (OIDC, client certificates, webhook). Kubernetes stores no user objects. Structured authentication configuration (GA 1.34) lets the API server trust several JWT issuers, with CEL claim mappings.
- Service accounts: identities for pods, managed by Kubernetes. Each namespace has a
defaultservice account. Create dedicated service accounts to give pods specific permissions.
Service Account Tokens
Pods get a short-lived, audience-bound projected token that the kubelet mounts and rotates automatically (the TokenRequest API). Since v1.24, long-lived Secret-based tokens are no longer auto-generated, and later releases clean up unused legacy tokens. Automounting is still on by default. Set automountServiceAccountToken: false on pods or service accounts that do not call the API.
Authentication Methods¶
| Method | Use Case |
|---|---|
| X.509 client certs | Break-glass admin access, kubeadm-generated kubeconfig |
| OIDC / structured JWT authentication (for example Okta, Entra ID, Keycloak) | Enterprise SSO for human users |
| Service account tokens (projected, bound) | Pod-to-API-server communication, workload identity federation |
| Webhook token auth | Custom token validation |
| Bootstrap tokens | Node bootstrapping (kubeadm) |
Pod certificates (GA 1.37) add a built-in way for pods to get X.509 certificates from a signer through PodCertificateRequest and a podCertificate projected volume, for mTLS without a sidecar CA. ClusterTrustBundle (GA 1.37) distributes CA bundles cluster-wide.
RBAC (Role-Based Access Control)¶
RBAC is the main authorization mechanism in Kubernetes. It uses four API objects:
| Object | Scope | Purpose |
|---|---|---|
Role |
Namespace | Defines permissions within a single namespace |
ClusterRole |
Cluster-wide | Defines permissions across all namespaces or on cluster-scoped resources |
RoleBinding |
Namespace | Grants a Role (or ClusterRole) to subjects within a namespace |
ClusterRoleBinding |
Cluster-wide | Grants a ClusterRole to subjects across all namespaces |
Subjects can be Users, Groups, or ServiceAccounts. The Node authorizer limits each kubelet to objects related to its own pods. Fine-grained kubelet API authorization (GA 1.36) splits the old all-or-nothing nodes/proxy permission into narrower subresources.
# Example: Read-only access to pods in a namespace
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
namespace: production
name: pod-reader
rules:
- apiGroups: [""]
resources: ["pods", "pods/log"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: read-pods
namespace: production
subjects:
- kind: User
name: developer
apiGroup: rbac.authorization.k8s.io
roleRef:
kind: Role
name: pod-reader
apiGroup: rbac.authorization.k8s.io
RBAC Best Practices
- Do not bind
cluster-adminbroadly. Use least-privilege Roles. create podsimplies access to every Secret and ServiceAccount in that namespace, because a pod can mount them.- Use namespace isolation to limit the blast radius.
Pod Security Standards and Admission¶
Pod Security Standards define three privilege levels, applied per namespace:
| Level | Description |
|---|---|
| Privileged | Unrestricted. Allows host access, privileged containers, and all capabilities |
| Baseline | Minimally restrictive. Blocks host namespaces, privileged containers, hostPath volumes, and added capabilities beyond a safe set |
| Restricted | Heavily restricted. Requires non-root, all capabilities dropped (except NET_BIND_SERVICE), no privilege escalation, and a RuntimeDefault/Localhost seccomp profile |
The built-in Pod Security Admission controller (GA since v1.25) enforces these levels using namespace labels. It has three modes per namespace:
- enforce: blocks pods that do not comply
- audit: records non-compliant pods in the audit log
- warn: returns warnings to the client without blocking
User namespaces (GA 1.36, hostUsers: false) map container root to an unprivileged host UID. That removes a whole class of container-escape risks.
NetworkPolicy¶
NetworkPolicies control traffic flow at the pod level:
- By default, all pod-to-pod traffic in a cluster is allowed
- A NetworkPolicy can restrict ingress and egress by pod or namespace label selectors, IP blocks, and ports
- The CNI plugin implements them (Calico, Cilium, Antrea, and others)
- With a CNI plugin that does not support NetworkPolicy, the rules are silently ignored
# Example: Deny all ingress to a namespace, allow only from specific pods
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: deny-all-ingress
namespace: production
spec:
podSelector: {}
policyTypes:
- Ingress
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-from-frontend
namespace: production
spec:
podSelector:
matchLabels:
app: api
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
app: frontend
ports:
- port: 8080
protocol: TCP
Default Deny
A common best practice is a default-deny NetworkPolicy in every namespace, with specific allow rules layered on top.
Secrets Management¶
Kubernetes Secrets¶
Kubernetes Secrets hold sensitive data (passwords, tokens, keys):
- Values are base64-encoded, which is not encryption. They are stored in etcd as plaintext unless you configure encryption at rest with an
EncryptionConfiguration(see How-to Guides) - KMS v2 (GA 1.29) is the recommended provider.
secretboxis a strong local alternative.aescbcis not recommended, and KMS v1 has been deprecated since v1.28 - Secrets are mounted into pods as files (tmpfs) or environment variables
- Secrets are namespace-scoped and cannot be read from another namespace without RBAC
External Secrets Management¶
For production, integrate with an external secret store:
- External Secrets Operator: syncs secrets from AWS Secrets Manager, Azure Key Vault, GCP Secret Manager, and HashiCorp Vault
- Sealed Secrets: encrypts secrets with a public key so they can live in Git, and decrypts them in-cluster with the private key
- Vault Agent Injector / Secrets Store CSI driver: mounts Vault or cloud secrets into pods as files
- SOPS: encrypts secret manifests for GitOps
Admission Controllers¶
Admission controllers intercept API requests after authentication and authorization, before the object is persisted:
| Controller | Purpose |
|---|---|
| PodSecurity | Enforces Pod Security Standards |
| ResourceQuota | Limits resource consumption per namespace |
| LimitRanger | Sets default resource requests and limits |
| NamespaceLifecycle | Prevents creation in terminating namespaces |
| ServiceAccount | Automates service account token mounting |
| NodeRestriction | Limits what a kubelet can modify (its own Node and its own pods) |
| ValidatingAdmissionPolicy / MutatingAdmissionPolicy | In-process CEL policies (GA 1.30 / 1.36). No webhook needed |
Policy Engines (OPA Gatekeeper / Kyverno)¶
External admission webhooks enforce custom policies:
- OPA Gatekeeper: uses Rego policies. Has an audit mode for dry runs and enforces through a webhook
- Kyverno: Kubernetes-native policies written in YAML, so no new language. Supports mutation and validation rules, and recent versions also generate CEL policies
- Both run as validating and mutating admission webhooks
- Policies can enforce image registry restrictions, required labels, resource limits, and security contexts
# Kyverno example: Require non-root containers
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-non-root
spec:
validationFailureAction: Enforce
rules:
- name: check-runAsNonRoot
match:
any:
- resources:
kinds:
- Pod
validate:
message: "Containers must run as non-root"
pattern:
spec:
containers:
- securityContext:
runAsNonRoot: true
Image Policy and Supply Chain¶
Image Policy Webhook¶
The ImagePolicyWebhook admission controller can check images against an external policy:
- Restrict registries (for example, allow only images from
registry.example.com) - Require signed images
- Enforce tag rules (for example, disallow
:latest)
Sigstore / Cosign¶
- Sign container images with Cosign (Sigstore)
- Verify signatures at admission time with policy controllers
- Kyverno and the Sigstore policy-controller both support Cosign verification. Gatekeeper supports it through external data providers
- Image volumes (GA 1.36) mount OCI artifacts as read-only volumes, for example model weights. Treat them as part of the supply chain too
Audit Logging¶
Kubernetes audit logging records API server requests for forensic analysis:
- Configured with
--audit-policy-file,--audit-log-path, or an audit webhook on the API server (see How-to Guides) - Four stages:
RequestReceived,ResponseStarted,ResponseComplete,Panic - Four levels:
None,Metadata,Request,RequestResponse - Essential for compliance (SOC 2, PCI-DSS) and incident investigation. Log Secrets at
Metadatalevel only
CIS Benchmarks¶
The Center for Internet Security publishes Kubernetes benchmarks that cover:
- Control plane configuration (API server, scheduler, controller-manager, etcd)
- Worker node configuration (kubelet, proxy)
- Network policies and CNI security
- RBAC and service account controls
- Secret encryption and audit logging
Automated scanning tools:
- kube-bench: open-source CIS benchmark scanner (Aqua Security)
- Trivy Operator: scans running workloads for misconfigurations and CVEs
Sources¶
- Kubernetes components
- Cluster architecture
- CRI documentation
- Cluster Networking
- Scheduling framework
- Dynamic Resource Allocation
- Sidecar containers
- Pod lifecycle
- Gateway API
- Kubernetes security overview
- Pod Security Standards
- Service accounts
- RBAC documentation
- Network Policies
- Encrypting Secret Data at Rest
- Audit logging
- CIS Kubernetes Benchmark
- KEP-1498: Kubernetes yearly support period
- KEP-2572: Defining the release cadence
- CHANGELOG-1.37