Skip to content

Explanation

What this page explains

Kubernetes has a declarative, state-driven architecture. A cluster has a control plane (the API server, etcd, the scheduler, and the controller managers) and worker nodes (kubelet, kube-proxy, and a CRI runtime). All state lives in etcd, and every component talks to the others only through the API server. This page covers why it works this way: the components, the reconciliation loop, pod lifecycle, networking, storage, scheduling, the 2025-2026 design shifts (DRA, in-place resize, cgroup v2, nftables, Gateway API), the release model, and the security model. Numbers and dates are in Reference. Tasks are in How-to Guides.

See also: Kubernetes hub, Reference, How-to Guides

Cluster Architecture Overview

The next diagram shows the control plane, two worker nodes, and the cluster DNS add-on. Only kube-apiserver talks to etcd.

graph TD
    subgraph ControlPlane["Control Plane"]
        API["kube-apiserver<br/>:6443"]
        ETCD["etcd<br/>:2379-2380"]
        SCHED["kube-scheduler<br/>:10259"]
        CM["kube-controller-manager<br/>:10257"]
        CCM["cloud-controller-manager<br/>(cloud only)"]
    end

    subgraph Addons["Cluster add-ons (run as pods)"]
        DNS["CoreDNS<br/>(Service kube-dns)"]
        CNIA["CNI agent DaemonSet<br/>(Cilium / Calico / Flannel)"]
    end

    subgraph Node1["Worker Node 1"]
        KUBELET1["kubelet<br/>:10250"]
        PROXY1["kube-proxy<br/>:10256"]
        CRI1["CRI runtime<br/>(containerd / CRI-O)"]
        POD1A["Pod A"]
        POD1B["Pod B"]
    end

    subgraph Node2["Worker Node 2"]
        KUBELET2["kubelet<br/>:10250"]
        PROXY2["kube-proxy<br/>:10256"]
        CRI2["CRI runtime<br/>(containerd / CRI-O)"]
        POD2A["Pod C"]
        POD2B["Pod D"]
    end

    API <--> ETCD
    SCHED -->|"watch + bind"| API
    CM -->|"watch + reconcile"| API
    CCM -->|"nodes, routes, LBs"| API
    DNS -->|"watch Services"| API
    KUBELET1 -->|"watch Pods, report status"| API
    PROXY1 -->|"watch EndpointSlices"| API
    KUBELET2 --> API
    PROXY2 --> API
    KUBELET1 -->|"CRI gRPC"| CRI1
    CRI1 --> POD1A
    CRI1 --> POD1B
    KUBELET2 -->|"CRI gRPC"| CRI2
    CRI2 --> POD2A
    CRI2 --> POD2B

Control Plane Components

kube-apiserver

The API server is the front end of the Kubernetes control plane. It exposes the Kubernetes HTTP API on port 6443 and is the only component that talks to etcd directly.

  • Validates and stores all API objects (pods, services, deployments, custom resources)
  • Serves the REST API used by kubectl, controllers, and external tools (JSON, Protobuf, and CBOR as an opt-in)
  • Runs authentication, authorization (RBAC, Node, webhook), and admission (built-in plugins, CEL ValidatingAdmissionPolicy / MutatingAdmissionPolicy, webhooks)
  • Scales horizontally: run several replicas behind a load balancer
  • Keeps a watch cache per resource, so the other components watch the API server, not etcd. Recent releases lower etcd load further: consistent reads from cache (GA 1.34), watch-cache initialization via streamed etcd range reads (EtcdRangeStream, beta 1.37), and gzip-compressed WatchList responses (beta 1.37)
  • API Priority and Fairness (APF) shares concurrency across tenants so one noisy client cannot starve the rest

etcd

etcd is a distributed, strongly consistent key-value store that holds all cluster state:

  • Stores every object definition, the cluster configuration, and dynamic state
  • Uses the Raft consensus algorithm for HA (a quorum of 3 or 5 members)
  • Listens on TCP 2379 (clients) and 2380 (peers)
  • Only the API server talks to etcd directly
  • Kubernetes v1.37 ships etcd 3.7.0 as its default version
  • Backups are critical. Take regular etcdctl snapshot save snapshots (see How-to Guides)

etcd Quorum

An etcd cluster needs a majority (quorum) to accept writes. A 3-member cluster survives 1 failure, and a 5-member cluster survives 2. Never run production etcd with fewer than 3 members. Even member counts add no fault tolerance.

kube-scheduler

The scheduler watches for new pods that have no node assigned and picks a node for each one. It is built on the scheduling framework, which is a set of plugin extension points:

  • PreFilter / Filter: remove nodes that cannot run the pod (resources, taints, node affinity, volume topology, DRA device availability)
  • PostFilter: preemption. If no node fits, lower-priority pods may be evicted
  • PreScore / Score: rank the remaining nodes with plugins such as NodeResourcesFit, NodeResourcesBalancedAllocation, InterPodAffinity, PodTopologySpread, TaintToleration, and ImageLocality
  • Reserve / Permit / PreBind / Bind: reserve resources (including DRA allocations) and write spec.nodeName through the pods/binding subresource
  • QueueingHints (GA 1.34) requeue unschedulable pods only when a relevant cluster event happens
  • Workload-aware scheduling (APIs beta in 1.37): the Workload and PodGroup APIs let the scheduler place a gang of pods all at once (for AI/ML training and MPI), with workload-aware preemption
  • Serves health and metrics on port 10259

kube-controller-manager

Runs the core reconciliation controllers in a single binary:

  • Node lifecycle controller: watches node heartbeats (Lease objects) and taints and evicts pods from unreachable nodes
  • ReplicaSet controller: keeps the desired number of pod replicas
  • Deployment controller: manages rolling updates and rollbacks through ReplicaSets
  • StatefulSet controller: ordered, identity-preserving pods. 1.37 adds an alpha Recreate update strategy
  • Job / CronJob controllers: batch completion, success policies, per-index backoff
  • EndpointSlice controller: publishes Service backends (the v1 Endpoints API is deprecated since 1.33)
  • HorizontalPodAutoscaler controller: scale to and from zero is on by default since 1.37
  • Service account and token controllers, Namespace controller (ordered namespace deletion is GA since 1.34), garbage collector

Each controller is a reconciliation loop. It watches the API server through shared informers and drives the current state toward the desired state.

cloud-controller-manager

This component holds the cloud-specific control loops: node initialization and addresses, routes, and type: LoadBalancer Services. They were split out of core Kubernetes so providers can release on their own schedule. In-tree cloud providers were removed in v1.31.

CoreDNS (cluster add-on)

CoreDNS provides cluster-wide DNS resolution for Services and Pods. It is not a control plane binary. It is an add-on that runs as a Deployment in kube-system behind a Service still named kube-dns for compatibility:

  • Resolves Service names to ClusterIPs (for example, my-service.my-namespace.svc.cluster.local)
  • Returns pod IPs directly for headless Services
  • Is configured through a Corefile (custom zones, forwarding, rewrites)
  • Has been the default since v1.13. The legacy kube-dns add-on is deprecated as of v1.37

Node Components

kubelet

The kubelet runs on every node and manages the pod lifecycle:

  • Watches the API server for pods assigned to its node
  • Pulls images and starts containers through the CRI runtime
  • Reports node and pod status back to the API server, plus a node Lease heartbeat
  • Runs liveness, readiness, and startup probes. 1.37 adds HTTP/2 cleartext (protocol field) probes
  • Mounts volumes (via CSI) and has the runtime set up pod networking (via CNI)
  • Applies in-place resizes of CPU and memory to running containers (GA 1.35)
  • Listens on port 10250 for the API server (exec, logs, port-forward, metrics)
  • Requires cgroup v2 by default since v1.35. It gets its cgroup driver from the runtime over CRI (KubeletCgroupDriverFromCRI, GA 1.34)

kube-proxy

kube-proxy programs each node's dataplane so that Service virtual IPs reach backend pods:

  • Watches Services and EndpointSlices
  • Modes on Linux: iptables (still the default), nftables (GA 1.33, the planned future default), and ipvs (deprecated since 1.35, with removal planned). The old userspace mode was removed in 1.26. Windows uses kernelspace
  • Some CNIs (for example Cilium's kube-proxy replacement) do the job of kube-proxy with eBPF instead
  • Serves a health check on port 10256

CRI Runtime (containerd / CRI-O)

The Container Runtime Interface (CRI) decouples the kubelet from any specific runtime:

  • containerd: the most common runtime and a CNCF graduated project. Kubernetes 1.36 and 1.37 are tested with containerd 2.x (2.2+/2.3+/2.4+). containerd 1.7 is outside the recommended matrix from 1.36 on
  • CRI-O: a lightweight runtime built only for Kubernetes. Its minor versions track Kubernetes minors
  • Both implement the CRI gRPC API that the kubelet calls
  • Image pulls, container creation, and execution are all delegated to the runtime
  • Dockershim was removed in Kubernetes v1.24. Images built with Docker still run, because they are OCI images

Container Network Interface (CNI)

CNI plugins configure pod networking:

  • Each pod gets its own IP address (IPv4, IPv6, or dual-stack)
  • Plugins include Cilium, Calico, Flannel, Antrea, the AWS VPC CNI, and GKE Dataplane V2
  • The plugin provides pod-to-pod connectivity within and across nodes
  • NetworkPolicy enforcement depends on the plugin (Flannel alone does not enforce it)

Container Storage Interface (CSI)

CSI is a standard interface for exposing storage to containers:

  • Out-of-tree CSI drivers have replaced the in-tree volume plugins. CSI migration is GA for the major cloud plugins, and in-tree code has been removed step by step (for example, Portworx operations are fully redirected to CSI as of v1.36)
  • Third-party drivers include AWS EBS CSI, GCE PD CSI, Azure Disk CSI, Ceph CSI, and Longhorn
  • Drivers handle provisioning, attachment, mounting, snapshots, cloning, expansion, and VolumeAttributesClass changes (GA 1.34)

Container Runtime Interface (CRI)

The CRI is the gRPC API between the kubelet and the container runtime:

  • Defines two services: RuntimeService (pod sandbox and container lifecycle) and ImageService (image management)
  • Lets runtimes plug in without changes to kubelet source code
  • Supports both container runtimes and sandboxed runtimes (Kata Containers, gVisor) through RuntimeClass
  • Keeps growing: 1.37 adds CheckpointPod / RestorePod RPCs for pod-level checkpoint and restore

Request Flow: Pod Creation

The next sequence diagram follows a bare Pod from kubectl apply to a running container.

sequenceDiagram
    actor User
    participant API as kube-apiserver
    participant etcd as etcd
    participant Sched as kube-scheduler
    participant Kubelet as kubelet
    participant CRI as CRI runtime

    User->>API: kubectl apply -f pod.yaml
    API->>API: authenticate + authorize (RBAC)
    API->>API: run admission (mutating, validating, CEL policies, webhooks)
    API->>etcd: persist pod spec (nodeName empty)
    etcd-->>API: confirmed
    API-->>User: 201 Created

    Sched->>API: watch for unscheduled pods
    API-->>Sched: pod with empty nodeName
    Sched->>Sched: filter + score nodes
    Sched->>API: POST pods/binding (set nodeName)
    API->>etcd: persist binding

    Kubelet->>API: watch pods for its node
    API-->>Kubelet: pod assigned to this node
    Kubelet->>CRI: RunPodSandbox, PullImage, CreateContainer, StartContainer
    CRI-->>Kubelet: container started
    Kubelet->>API: update pod status (Running)
    API->>etcd: persist status

How It Works

Desired-state reconciliation, pod lifecycle, networking, storage, and scheduling internals.

Desired-State Reconciliation Loop

The next sequence diagram shows how a Deployment becomes pods through a chain of independent controllers. Each controller only reacts to watched changes.

sequenceDiagram
    participant User as User / CI
    participant API as kube-apiserver
    participant ETCD as etcd
    participant Ctrl as kube-controller-manager
    participant Sched as kube-scheduler
    participant KL as kubelet (Node)
    participant CRI as containerd

    User->>API: kubectl apply -f deployment.yaml
    API->>ETCD: Store desired state
    Ctrl->>API: Watch: new Deployment
    Ctrl->>API: Deployment controller creates ReplicaSet
    Ctrl->>API: ReplicaSet controller creates Pods
    Sched->>API: Watch: unscheduled Pods
    Sched->>Sched: Filter and score nodes (resources, affinity, taints)
    Sched->>API: Bind Pod to Node
    KL->>API: Watch: Pod assigned to my node
    KL->>CRI: Create pod sandbox
    CRI->>CRI: Pull image, start containers
    KL->>API: Update Pod status: Running

    loop Reconciliation
        Ctrl->>API: Watch: actual vs desired state
        Ctrl->>Ctrl: If replicas below desired, create Pods
        Ctrl->>Ctrl: If replicas above desired, delete surplus
    end

Pod Lifecycle

The next state diagram shows the pod phases. Init containers, including native sidecars, run while the pod is Pending. Readiness is a condition inside Running, not a phase.

stateDiagram-v2
    [*] --> Pending: Pod accepted
    state Pending {
        [*] --> Unscheduled
        Unscheduled --> Scheduled: bound to node
        Scheduled --> InitContainers: image pull, sandbox
        InitContainers --> Starting: init done, sidecars running
    }
    Pending --> Running: all containers started
    Running --> Succeeded: all containers exit 0 (restartPolicy Never or OnFailure)
    Running --> Failed: container exits non-zero, not restarted
    Running --> Unknown: node unreachable
    Unknown --> Running: node recovers
    Unknown --> Failed: node lost, pod evicted
    Succeeded --> [*]
    Failed --> [*]

    state Running {
        [*] --> NotReady
        NotReady --> Ready: readiness probe passes
        Ready --> NotReady: readiness probe fails
    }

Native sidecars (GA 1.33) are init containers with restartPolicy: Always. They start before the app containers, keep running for the pod's lifetime, and stop after the app containers. This fixes the old problems of Jobs never finishing and of shutdown ordering. In-place resize (GA 1.35) changes a running container's CPU and memory through the resize subresource. The kubelet reports progress with the PodResizePending and PodResizeInProgress conditions, and the pod keeps running.

Networking Model

The 4 Networking Rules

  1. Pod-to-Pod: every Pod gets its own IP. All Pods can talk to each other without NAT.
  2. Pod-to-Service: Services provide stable virtual IPs (ClusterIP), implemented by kube-proxy (iptables, nftables, ipvs) or an eBPF replacement.
  3. External-to-Service: LoadBalancer, NodePort, or Ingress / Gateway API expose Services.
  4. Pod-to-External: Pods reach external networks through SNAT.

The next diagram shows those paths for a Service backed by pods on two nodes.

flowchart TB
    subgraph Cluster["Kubernetes Cluster"]
        subgraph Node1["Node 1"]
            P1["Pod A<br/>10.244.1.2"]
            P2["Pod B<br/>10.244.1.3"]
            KP1["kube-proxy<br/>(iptables / nftables)"]
        end

        subgraph Node2["Node 2"]
            P3["Pod C<br/>10.244.2.2"]
            P4["Pod D<br/>10.244.2.3"]
            KP2["kube-proxy"]
        end

        SVC["Service my-svc<br/>ClusterIP 10.96.0.10<br/>backends: Pod A, Pod C"]
        GW["Gateway + HTTPRoute<br/>(Gateway API controller)"]
        CNI["CNI plugin<br/>(Cilium / Calico / Flannel)<br/>pod-to-pod routing"]
    end

    External["External traffic"] -->|"LoadBalancer / NodePort"| GW
    GW --> SVC
    KP1 -.->|"programs rules"| SVC
    KP2 -.->|"programs rules"| SVC
    SVC -->|"DNAT"| P1
    SVC -->|"DNAT"| P3
    P1 <-->|"CNI overlay or native routing"| P3
    P2 <-->|"CNI"| P4
    CNI -.-> P1
    CNI -.-> P3

Ingress vs Gateway API

Ingress (networking.k8s.io/v1) is frozen: it still works, but new work happens in Gateway API. Gateway API is a set of role-oriented CRDs: the infrastructure provider owns GatewayClass, the cluster operator owns Gateway, and application developers own HTTPRoute, GRPCRoute, TLSRoute, TCPRoute, and UDPRoute. It supports traffic splitting, header matching, and cross-namespace delegation with ReferenceGrant. The widely used Ingress NGINX controller was retired in March 2026. That made Gateway API the default path for new clusters. See How-to Guides.

Storage Architecture

The next diagram shows dynamic provisioning from a PVC down to a real disk.

flowchart LR
    Pod["Pod"] --> PVC["PersistentVolumeClaim<br/>(request: 10Gi)"]
    PVC --> PV["PersistentVolume<br/>(10Gi, RWO)"]
    PVC -.->|"storageClassName"| SC["StorageClass<br/>(provisioner: ebs.csi.aws.com)"]
    SC --> CSI["CSI driver<br/>(controller + node plugins)"]
    CSI -->|"CreateVolume"| PV
    CSI --> Disk["Cloud disk or<br/>storage backend"]

Scheduling Algorithm

Phase Operation
Filtering Remove nodes that do not meet the pod's requirements: resources (NodeResourcesFit), taints, node affinity, volume topology, DRA devices
Scoring Rank the remaining nodes. The default plugins include NodeResourcesFit (LeastAllocated strategy), NodeResourcesBalancedAllocation, NodeAffinity, InterPodAffinity, PodTopologySpread, TaintToleration, and ImageLocality
Binding Assign the pod to the highest-scoring node through the binding subresource
Preemption PostFilter: if no node fits, evict lower-priority pods. 1.37 adds alpha preemption for deferred in-place resizes

Dynamic Resource Allocation

DRA (core GA in v1.34, API resource.k8s.io/v1) replaces the device-plugin model, which could only count opaque integers such as nvidia.com/gpu: 1, with claims for devices that have attributes. Drivers publish ResourceSlice objects that describe devices and their attributes. Admins define DeviceClass objects. Workloads request devices through a ResourceClaim or a ResourceClaimTemplate with CEL selectors. The scheduler allocates matching devices itself, so it knows about devices before it binds the pod. Releases since GA have added prioritized alternatives and admin access (GA 1.36), and device taints/tolerations and extended-resource mapping (GA 1.37). That second feature lets existing nvidia.com/gpu-style requests be served by DRA drivers.

The next sequence diagram shows the DRA allocation flow for a pod that references a claim template.

sequenceDiagram
    participant Drv as DRA driver (kubelet plugin)
    participant API as kube-apiserver
    participant RCC as resourceclaim controller
    participant Sched as kube-scheduler
    participant KL as kubelet

    Drv->>API: publish ResourceSlice (devices + attributes)
    Note over API: Pod references ResourceClaimTemplate
    RCC->>API: create ResourceClaim for the Pod
    Sched->>API: watch Pod, ResourceClaim, ResourceSlices, DeviceClasses
    Sched->>Sched: filter nodes by CEL selectors and free devices
    Sched->>API: write allocation to ResourceClaim status
    Sched->>API: bind Pod to node
    KL->>Drv: NodePrepareResources (gRPC)
    Drv-->>KL: CDI device IDs
    KL->>KL: start containers with CDI devices injected
    KL->>Drv: NodeUnprepareResources when the Pod ends

Design Shifts in 2025-2026

Shift What changed Why
Devices as first-class API DRA core GA (1.34). Feature set filled in through 1.37 GPUs, NICs, and accelerators need attributes, sharing, and topology that device plugins cannot express
Mutable pod resources In-place resize GA (1.35), including init containers (GA 1.37) and pod-level resources (beta) Vertical scaling without restarts, for JVMs, AI inference, and VPA
Lifecycle-correct sidecars Native sidecars GA (1.33) Service mesh and log-agent containers need defined start and stop ordering
Linux baseline raised cgroup v1 refused by default (1.35). User namespaces GA (1.36). Rootless kubelet beta (1.37) cgroup v2 is what PSI metrics, MemoryQoS, and swap need. User namespaces shrink the container-escape blast radius
Dataplane modernisation nftables GA (1.33). ipvs deprecated (1.35). Default-mode switch announced (1.37) iptables rule-update cost grows with Service count. ipvs was under-maintained
North-south traffic Gateway API replaces Ingress. Ingress NGINX retired (March 2026) Ingress was under-specified (annotation sprawl), and the controller lacked maintainers
In-tree policy CEL ValidatingAdmissionPolicy (GA 1.30) and MutatingAdmissionPolicy (GA 1.36) Webhooks add latency and failure modes. CEL runs inside the API server
AI/ML batch Workload-aware / gang scheduling APIs (beta 1.37). HPA scale-to-zero on by default (1.37) Distributed training needs all-or-nothing placement. Idle GPU inference should scale to zero

Release Cadence and Support Model

Kubernetes releases about three minor versions a year, roughly every 15-17 weeks (1.34 in August 2025, 1.35 in December 2025, 1.36 in April 2026, 1.37 in August 2026). The project supports the latest three minors. Each gets about 12 months of patches plus a 2-month maintenance-mode tail, so any version lives about 14 months. Two KEPs set this model: KEP-1498 gave each release a one-year support period (from v1.19), and KEP-2572 moved Kubernetes from four releases a year to three (from 2021). The skew policy lets kubelets lag the API server by three minors. That means node fleets can be upgraded less often than control planes. Control planes still have to step through every minor, one at a time. Managed providers add paid extended support (EKS, GKE Extended channel, AKS LTS) for organisations that cannot keep up with that cadence. Current dates are in Reference.

Scalability and Performance

Upstream does not publish "maximums". It publishes a tested envelope: 5,000 nodes, 150,000 pods, and at most 110 pods per node. Inside that envelope, the SIG Scalability SLOs hold, for example p99 mutating API latency of 1 s or less, and p99 stateless pod startup of 5 s or less. Going beyond the envelope is possible but becomes your own engineering problem. The limiting factors are usually etcd write latency and database size, watch fan-out from the API server, and scheduler throughput with complex affinity. Hyperscalers have gone far past the envelope by replacing parts of the stack: GKE supports 65,000-node clusters, and in 2025 it ran an experimental 130,000-node cluster on a Spanner-based store. The thresholds, SLOs, and indicative figures are in Reference.

Security Model

Kubernetes security follows the 4C's model: Cloud, Cluster, Container, and Code. This section covers the cluster and container layers: identity, access control, pod isolation, secrets, and supply chain. The setup tasks are in How-to Guides, and the hardening checklist is in Reference.

Authentication and Identity

Users and Service Accounts

Kubernetes has two kinds of identity:

  • Users: managed outside the cluster (OIDC, client certificates, webhook). Kubernetes stores no user objects. Structured authentication configuration (GA 1.34) lets the API server trust several JWT issuers, with CEL claim mappings.
  • Service accounts: identities for pods, managed by Kubernetes. Each namespace has a default service account. Create dedicated service accounts to give pods specific permissions.

Service Account Tokens

Pods get a short-lived, audience-bound projected token that the kubelet mounts and rotates automatically (the TokenRequest API). Since v1.24, long-lived Secret-based tokens are no longer auto-generated, and later releases clean up unused legacy tokens. Automounting is still on by default. Set automountServiceAccountToken: false on pods or service accounts that do not call the API.

Authentication Methods

Method Use Case
X.509 client certs Break-glass admin access, kubeadm-generated kubeconfig
OIDC / structured JWT authentication (for example Okta, Entra ID, Keycloak) Enterprise SSO for human users
Service account tokens (projected, bound) Pod-to-API-server communication, workload identity federation
Webhook token auth Custom token validation
Bootstrap tokens Node bootstrapping (kubeadm)

Pod certificates (GA 1.37) add a built-in way for pods to get X.509 certificates from a signer through PodCertificateRequest and a podCertificate projected volume, for mTLS without a sidecar CA. ClusterTrustBundle (GA 1.37) distributes CA bundles cluster-wide.

RBAC (Role-Based Access Control)

RBAC is the main authorization mechanism in Kubernetes. It uses four API objects:

Object Scope Purpose
Role Namespace Defines permissions within a single namespace
ClusterRole Cluster-wide Defines permissions across all namespaces or on cluster-scoped resources
RoleBinding Namespace Grants a Role (or ClusterRole) to subjects within a namespace
ClusterRoleBinding Cluster-wide Grants a ClusterRole to subjects across all namespaces

Subjects can be Users, Groups, or ServiceAccounts. The Node authorizer limits each kubelet to objects related to its own pods. Fine-grained kubelet API authorization (GA 1.36) splits the old all-or-nothing nodes/proxy permission into narrower subresources.

# Example: Read-only access to pods in a namespace
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  namespace: production
  name: pod-reader
rules:
- apiGroups: [""]
  resources: ["pods", "pods/log"]
  verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: read-pods
  namespace: production
subjects:
- kind: User
  name: developer
  apiGroup: rbac.authorization.k8s.io
roleRef:
  kind: Role
  name: pod-reader
  apiGroup: rbac.authorization.k8s.io

RBAC Best Practices

  • Do not bind cluster-admin broadly. Use least-privilege Roles.
  • create pods implies access to every Secret and ServiceAccount in that namespace, because a pod can mount them.
  • Use namespace isolation to limit the blast radius.

Pod Security Standards and Admission

Pod Security Standards define three privilege levels, applied per namespace:

Level Description
Privileged Unrestricted. Allows host access, privileged containers, and all capabilities
Baseline Minimally restrictive. Blocks host namespaces, privileged containers, hostPath volumes, and added capabilities beyond a safe set
Restricted Heavily restricted. Requires non-root, all capabilities dropped (except NET_BIND_SERVICE), no privilege escalation, and a RuntimeDefault/Localhost seccomp profile

The built-in Pod Security Admission controller (GA since v1.25) enforces these levels using namespace labels. It has three modes per namespace:

  • enforce: blocks pods that do not comply
  • audit: records non-compliant pods in the audit log
  • warn: returns warnings to the client without blocking

User namespaces (GA 1.36, hostUsers: false) map container root to an unprivileged host UID. That removes a whole class of container-escape risks.

NetworkPolicy

NetworkPolicies control traffic flow at the pod level:

  • By default, all pod-to-pod traffic in a cluster is allowed
  • A NetworkPolicy can restrict ingress and egress by pod or namespace label selectors, IP blocks, and ports
  • The CNI plugin implements them (Calico, Cilium, Antrea, and others)
  • With a CNI plugin that does not support NetworkPolicy, the rules are silently ignored
# Example: Deny all ingress to a namespace, allow only from specific pods
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: deny-all-ingress
  namespace: production
spec:
  podSelector: {}
  policyTypes:
  - Ingress
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-from-frontend
  namespace: production
spec:
  podSelector:
    matchLabels:
      app: api
  policyTypes:
  - Ingress
  ingress:
  - from:
    - podSelector:
        matchLabels:
          app: frontend
    ports:
    - port: 8080
      protocol: TCP

Default Deny

A common best practice is a default-deny NetworkPolicy in every namespace, with specific allow rules layered on top.

Secrets Management

Kubernetes Secrets

Kubernetes Secrets hold sensitive data (passwords, tokens, keys):

  • Values are base64-encoded, which is not encryption. They are stored in etcd as plaintext unless you configure encryption at rest with an EncryptionConfiguration (see How-to Guides)
  • KMS v2 (GA 1.29) is the recommended provider. secretbox is a strong local alternative. aescbc is not recommended, and KMS v1 has been deprecated since v1.28
  • Secrets are mounted into pods as files (tmpfs) or environment variables
  • Secrets are namespace-scoped and cannot be read from another namespace without RBAC

External Secrets Management

For production, integrate with an external secret store:

  • External Secrets Operator: syncs secrets from AWS Secrets Manager, Azure Key Vault, GCP Secret Manager, and HashiCorp Vault
  • Sealed Secrets: encrypts secrets with a public key so they can live in Git, and decrypts them in-cluster with the private key
  • Vault Agent Injector / Secrets Store CSI driver: mounts Vault or cloud secrets into pods as files
  • SOPS: encrypts secret manifests for GitOps

Admission Controllers

Admission controllers intercept API requests after authentication and authorization, before the object is persisted:

Controller Purpose
PodSecurity Enforces Pod Security Standards
ResourceQuota Limits resource consumption per namespace
LimitRanger Sets default resource requests and limits
NamespaceLifecycle Prevents creation in terminating namespaces
ServiceAccount Automates service account token mounting
NodeRestriction Limits what a kubelet can modify (its own Node and its own pods)
ValidatingAdmissionPolicy / MutatingAdmissionPolicy In-process CEL policies (GA 1.30 / 1.36). No webhook needed

Policy Engines (OPA Gatekeeper / Kyverno)

External admission webhooks enforce custom policies:

  • OPA Gatekeeper: uses Rego policies. Has an audit mode for dry runs and enforces through a webhook
  • Kyverno: Kubernetes-native policies written in YAML, so no new language. Supports mutation and validation rules, and recent versions also generate CEL policies
  • Both run as validating and mutating admission webhooks
  • Policies can enforce image registry restrictions, required labels, resource limits, and security contexts
# Kyverno example: Require non-root containers
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
  name: require-non-root
spec:
  validationFailureAction: Enforce
  rules:
  - name: check-runAsNonRoot
    match:
      any:
      - resources:
          kinds:
          - Pod
    validate:
      message: "Containers must run as non-root"
      pattern:
        spec:
          containers:
          - securityContext:
              runAsNonRoot: true

Image Policy and Supply Chain

Image Policy Webhook

The ImagePolicyWebhook admission controller can check images against an external policy:

  • Restrict registries (for example, allow only images from registry.example.com)
  • Require signed images
  • Enforce tag rules (for example, disallow :latest)

Sigstore / Cosign

  • Sign container images with Cosign (Sigstore)
  • Verify signatures at admission time with policy controllers
  • Kyverno and the Sigstore policy-controller both support Cosign verification. Gatekeeper supports it through external data providers
  • Image volumes (GA 1.36) mount OCI artifacts as read-only volumes, for example model weights. Treat them as part of the supply chain too

Audit Logging

Kubernetes audit logging records API server requests for forensic analysis:

  • Configured with --audit-policy-file, --audit-log-path, or an audit webhook on the API server (see How-to Guides)
  • Four stages: RequestReceived, ResponseStarted, ResponseComplete, Panic
  • Four levels: None, Metadata, Request, RequestResponse
  • Essential for compliance (SOC 2, PCI-DSS) and incident investigation. Log Secrets at Metadata level only

CIS Benchmarks

The Center for Internet Security publishes Kubernetes benchmarks that cover:

  • Control plane configuration (API server, scheduler, controller-manager, etcd)
  • Worker node configuration (kubelet, proxy)
  • Network policies and CNI security
  • RBAC and service account controls
  • Secret encryption and audit logging

Automated scanning tools:

  • kube-bench: open-source CIS benchmark scanner (Aqua Security)
  • Trivy Operator: scans running workloads for misconfigurations and CVEs

Sources