Skip to content

Kubernetes

Summary

Kubernetes (K8s) is the industry-standard, open-source container orchestrator. You declare the desired state of your workloads, and control loops keep reconciling the cluster toward it: scheduling, scaling, service discovery, storage, and self-healing, across fleets of machines. It is Apache-2.0 licensed and was the first project to graduate from the CNCF (2018). The latest minor is v1.37 "Garhwal" (2026-08-26), with patch v1.37.1. Upstream supports 1.37, 1.36, and 1.35, and 1.34 reaches end of life on 2026-10-27. Changes in 2025-2026: DRA, in-place pod resize, and native sidecars went GA. cgroup v1 and kube-proxy ipvs mode are on their way out. Ingress NGINX was retired in favour of Gateway API.

Key Facts

Attribute Detail
Latest Version v1.37.1 (2026-09-15, per the 1.37 changelog and patch schedule). Minor v1.37.0 "Garhwal" released 2026-08-26
Supported branches 1.37 (EOL 2027-10-28), 1.36 (EOL 2027-06-28), 1.35 (EOL 2027-02-28). 1.34 is in maintenance mode, EOL 2026-10-27
Release cadence About 3 minor releases a year. Each gets ~12 months of patches plus 2 months of maintenance mode. Monthly patch releases
Repository github.com/kubernetes/kubernetes
Stars ~115k+ (as of 2026-07)
Language Go
License Apache 2.0
Governance CNCF (Linux Foundation). First CNCF graduated project (March 2018). Community-run through SIGs and a Steering Committee
First release v1.0, July 2015 (originally designed at Google, drawing on Borg)
Default runtime / etcd containerd 2.x (or CRI-O). v1.37 bundles etcd 3.7.0 and CoreDNS 1.14.6

Full release, EOL, and skew tables are in Reference.

Overview

Kubernetes is a production-grade container orchestration system. It was originally developed at Google and is now maintained by the Kubernetes community under the Cloud Native Computing Foundation (CNCF). It uses a desired-state model: controllers continuously reconcile the actual state with the declared intent. Kubernetes manages the whole lifecycle of containerized applications: scheduling, scaling, networking, storage, and self-healing. Its API is extensible through CustomResourceDefinitions and operators, which makes it a platform for building platforms.

What Changed in 2025-2026

Release Headline
v1.33 "Octarine" (2025-04) Native sidecar containers GA. kube-proxy nftables GA. v1 Endpoints deprecated
v1.34 "Of Wind & Will" (2025-08) DRA core GA (resource.k8s.io/v1). Structured authentication config GA. VolumeAttributesClass GA
v1.35 "Timbernetes" (2025-12) In-place pod resize GA. kubelet refuses cgroup v1 by default. ipvs mode deprecated
v1.36 "Haru" (2026-04) User namespaces GA. MutatingAdmissionPolicy GA. ImageVolume GA. gitRepo volumes disabled
v1.37 "Garhwal" (2026-08) SELinuxMount GA (action required). Pod certificates and ClusterTrustBundle GA. metrics.k8s.io/v1. KYAML stable. HPA scale-to-zero on by default. Rootless kubelet beta. Workload-aware scheduling APIs beta. kube-dns deprecated
Ecosystem Ingress NGINX retired (March 2026). Gateway API v1.6 (TCPRoute/UDPRoute GA). containerd 2.3 LTS and 2.4

Details and the full removals list: Reference: Feature Graduations by Release and Deprecations and Removals. For upgrade blockers, see How-to: Prepare for the v1.37 Upgrade.

Evaluation

  • Why it is better: cloud-agnostic, with a huge ecosystem (the CNCF landscape) and a declarative, extensible API. Built-in self-healing, autoscaling (HPA, VPA, Cluster Autoscaler/Karpenter), service discovery, and rolling updates. It is the dominant industry standard, and every major cloud offers it as a managed service.

  • When it fits (Applicability):

  • Microservices at scale across many nodes
  • CI/CD with automated rollouts and rollbacks, including GitOps (Argo CD, Flux)
  • Multi-cloud / hybrid cloud portability
  • Stateful workloads with persistent volumes and operators
  • AI/ML training and inference (DRA for GPUs, gang scheduling, scale-to-zero)
  • Edge deployments (K3s, MicroK8s, Talos)

  • When it does not fit: a handful of containers on one host (use Docker Compose or Podman), teams without platform-engineering capacity that cannot use a managed service, and pure VM estates (consider OpenStack, Proxmox, or KubeVirt as a bridge).

  • Pros and Cons:

Pros Cons
Cloud-agnostic, runs anywhere Steep learning curve
Self-healing, autoscaling Complex networking (CNI plugins, Gateway API controllers)
Massive CNCF ecosystem Control plane overhead for small workloads
Declarative desired-state model, extensible via CRDs YAML verbosity
Gateway API, service mesh integrations Security hardening requires expertise
GPU/DRA scheduling for AI/ML etcd operational complexity
Every major cloud offers managed K8s Fast cadence: about 14 months of support per minor means yearly upgrades are mandatory

Architecture

The compact diagram below shows the control plane and one worker node. The full component breakdown and request flows are in Explanation.

flowchart TB
    subgraph ControlPlane["Control Plane"]
        API["kube-apiserver<br/>(REST API, admission)"]
        ETCD["etcd<br/>(Raft KV store)"]
        Sched["kube-scheduler<br/>(pod placement)"]
        CM["kube-controller-manager<br/>(reconciliation loops)"]
        CCM["cloud-controller-manager<br/>(cloud API integration)"]
    end

    subgraph WorkerNode["Worker Node"]
        Kubelet["kubelet<br/>(pod lifecycle)"]
        KProxy["kube-proxy<br/>(iptables / nftables)"]
        CRI["containerd / CRI-O"]
        Pods["Pods"]
    end

    API <-->|"read/write state"| ETCD
    Sched -->|"watch, bind pod"| API
    CM -->|"watch, reconcile"| API
    CCM -->|"nodes, LBs, routes"| API
    Kubelet -->|"watch pods, report status"| API
    KProxy -->|"watch EndpointSlices"| API
    Kubelet -->|"CRI gRPC"| CRI
    CRI --> Pods
    KProxy -.->|"Service rules"| Pods

Key Features

Feature Detail
Pod Scheduling Affinity, anti-affinity, taints, tolerations, topology spread, gang scheduling (beta 1.37)
Auto-Scaling HPA (scale-to-zero on by default since 1.37), VPA with in-place resize (GA 1.35), Cluster Autoscaler / Karpenter
Service Discovery ClusterIP, NodePort, LoadBalancer, ExternalName, headless. EndpointSlices
Gateway API / Ingress L4/L7 routing and TLS. Gateway API is the successor to Ingress
Storage PV, PVC, CSI drivers, StorageClasses, VolumeAttributesClass, snapshots
ConfigMaps / Secrets Externalized configuration and credentials. Encryption at rest via KMS v2
RBAC Fine-grained role-based access control. CEL admission policies
Namespaces Logical cluster partitioning
DRA Dynamic Resource Allocation for GPUs, NICs, and FPGAs. Core GA since v1.34
Sidecar containers Native sidecars (init containers with restartPolicy: Always). GA since v1.33
Custom Resources Extend the API with CRDs + Operators

Key Ecosystem

Category Tools
Managed K8s EKS, GKE, AKS, DOKS, OKE, Linode LKE
Lightweight / distros K3s, MicroK8s, Kind, Minikube, Talos, OpenShift, Rancher RKE2
Networking Cilium, Calico, Flannel, Antrea
Gateway API controllers Envoy Gateway, Istio, Cilium, NGINX Gateway Fabric, Kong, cloud LBs
Service Mesh Istio, Linkerd, Consul Connect
GitOps Argo CD, Flux
Observability Prometheus, Grafana, OpenTelemetry, VictoriaMetrics
Security Falco, OPA/Gatekeeper, Kyverno, Trivy, kube-bench

Pricing

Offering Cost (as of 2026-09) Notes
Self-hosted Free (Apache 2.0) You manage everything
AWS EKS $0.10/hr per cluster (standard support), $0.60/hr (extended support), plus node costs Managed control plane
Google GKE $0.10/hr per cluster management fee (free-tier credit covers one zonal or Autopilot cluster). Autopilot bills pod requests instead of nodes Extended channel costs extra
Azure AKS Free tier (no SLA). Standard tier $0.10/hr. Premium $0.60/hr (includes LTS) Nodes billed separately
Enterprise Various (OpenShift, Rancher Prime, Tanzu) Support + add-ons

Pricing volatility

Cloud prices change. Confirm on each provider's pricing page. The shared $0.60/hr extended-support rate is a strong incentive to stay inside the upstream support window.

Compatibility

Dimension Support
Container runtimes containerd (2.3+/2.4+ recommended for 1.37), CRI-O (matching minor). Any CRI v1 runtime
Node OS Linux with cgroup v2 (required by default since 1.35). Windows Server worker nodes
CPU architecture amd64, arm64, ppc64le, s390x (official server binaries)
Storage CSI (Ceph, EBS, GCE PD, Azure Disk, NFS, Longhorn, and others)
Networking CNI plugins (Cilium, Calico, Flannel, and others). IPv4, IPv6, dual-stack
Infrastructure Bare metal, VMs, any cloud, edge

Scale Limits (Upstream)

The upstream-tested envelope is 5,000 nodes, 150,000 pods, 300,000 containers, and at most 110 pods per node, with SLOs such as p99 mutating API latency of 1 s or less. Managed offerings go further: GKE supports 65,000 nodes and EKS up to 100,000. The full threshold and SLO tables are in Reference: Scalability Thresholds.

Alternatives and Lock-in

  • Alternatives: HashiCorp Nomad (simpler scheduler), Docker Swarm (legacy), AWS ECS / Google Cloud Run / Azure Container Apps (proprietary serverless containers), and OpenStack or OpenNebula for VM-first estates.
  • Lock-in: the upstream API is portable and certified by the CNCF Certified Kubernetes conformance programme. Lock-in comes mostly from cloud-specific add-ons (IAM integration, load balancers, storage classes, CNI) and from the operator ecosystem. Tools such as Velero, GitOps repositories, and Cluster API keep clusters reproducible.

Topic Map

  • How-to Guides: plan a production cluster, upgrade kubeadm to v1.37, move off Ingress NGINX, tune, secure, monitor, back up, kubectl recipes.
  • Reference: release branches and EOL dates, version skew, GA graduations and removals (v1.33 to v1.37), runtime compatibility, ports, scalability limits, hardening checklist.
  • Explanation: control plane and node components, control plane and node components, CNI/CSI/CRI, the pod-creation request flow, DRA, 2025-2026 design shifts, release model, scalability, security model.

Comparisons:

Sources

Source URL Retrieved Via
Official Website https://kubernetes.io Direct
Documentation https://kubernetes.io/docs/ Direct
GitHub Repository https://github.com/kubernetes/kubernetes Direct
Releases https://kubernetes.io/releases/ Web Search
Release schedule data https://github.com/kubernetes/website/blob/main/data/releases/schedule.yaml raw.githubusercontent.com
CHANGELOG-1.37 https://github.com/kubernetes/kubernetes/blob/master/CHANGELOG/CHANGELOG-1.37.md raw.githubusercontent.com
v1.37 release blog https://kubernetes.io/blog/2026/08/26/kubernetes-v1-37-release/ Web Search
v1.36 release blog https://kubernetes.io/blog/2026/04/22/kubernetes-v1-36-release/ Web Search
Release team info https://www.kubernetes.dev/resources/release/ Web Search
Version Skew Policy https://kubernetes.io/releases/version-skew-policy/ raw.githubusercontent.com (website source)
API Reference https://kubernetes.io/docs/reference/kubernetes-api/ Direct
CNCF https://cncf.io Direct
CNCF Landscape https://landscape.cncf.io Direct
Scalability Thresholds https://github.com/kubernetes/community/blob/master/sig-scalability/configs-and-limits/thresholds.md raw.githubusercontent.com
Ingress NGINX Retirement https://kubernetes.io/blog/2025/11/11/ingress-nginx-retirement/ Web Search + repo README
containerd releases / K8s matrix https://github.com/containerd/containerd/blob/main/RELEASES.md raw.githubusercontent.com
Gateway API https://gateway-api.sigs.k8s.io/ Direct
The Register on v1.37 removals https://www.theregister.com/devops/2026/08/26/kubernetes-cleans-house-bins-legacy-kube-dns-ipvs-and-cgroup-v1/5292717 Web Search
GKE 65k-node clusters https://cloud.google.com/blog/products/containers-kubernetes/gke-65k-nodes-and-counting Web Search
EKS extended support pricing https://aws.amazon.com/about-aws/whats-new/2024/04/amazon-eks-support-kubernetes-versions/ Web Search

Questions

Open Questions

  • Does v1.38 (planned for 2026-12-16 per the SIG Release schedule) finally remove the containerd 1.7 kubelet fallbacks (deferred from 1.36 and 1.37)?
  • Will nftables become the kube-proxy default on Linux in 1.38 or later, and which kernels will it need?
  • The ipvs (disable in 1.40, remove in 1.43) and kube-dns (last build 1.40) dates come from release coverage. Confirm them against the KEPs.
  • How mature are the DRA drivers (NVIDIA, AMD, Intel) for production GPU partitioning now that DRA extended-resource mapping is GA?

Answered Questions

  • How does SELinuxMount affect pod startup in SELinux-enforcing environments? SELinuxMount (KEP-1710) mounts a volume with the right label via -o context=..., instead of recursively relabelling every file. On large volumes this removes the per-file setxattr cost at pod start. The narrower SELinuxMountReadWriteOncePod variant came first. The general SELinuxMount gate only went GA and on by default in v1.37, as an "ACTION REQUIRED" change: pods that share a volume with different SELinux labels can break, so check them on v1.36 first or set seLinuxChangePolicy: Recursive. No upstream benchmark figures were found. Resolved via CHANGELOG-1.37.
  • What is the production readiness of DRA for GPU partitioning? DRA core went GA in v1.34 (resource.k8s.io/v1). Its feature gate has been locked on since 1.35. Prioritized lists and admin access went GA in 1.36, and device taints and extended-resource mapping went GA in 1.37. Partitionable devices and consumable capacity are still maturing. Driver maturity is vendor-specific (see Open Questions). Resolved via CHANGELOG-1.34 to 1.37.
  • How does Gateway API adoption compare to the retired ingress-nginx in production? Gateway API is the official successor to Ingress. ingress-nginx entered best-effort maintenance in November 2025 and was retired in March 2026, with no further releases or security fixes. Production-grade controllers include Envoy Gateway, Istio, Cilium, NGINX Gateway Fabric, and Kong. Gateway API adds traffic splitting, header matching, and weighted backends, and uses role-oriented resources (GatewayClass for infrastructure, Gateway for operators, Routes for app teams). ingress2gateway converts ingress-nginx objects. Snippet-heavy configurations need a redesign. Resolved via the ingress-nginx README and the Gateway API docs.
  • What are the recommended etcd backup strategies for clusters with >10,000 objects? (1) Take periodic etcdctl snapshot save snapshots every 6-12 hours and before every upgrade. (2) Rely on the API server's periodic compaction (every 5 min by default) and defragment members one at a time in maintenance windows. (3) Store snapshots in external object storage (S3/GCS). (4) Keep several generations. (5) Test etcdutl snapshot restore regularly on a separate cluster. Resolved via the etcd and Kubernetes documentation.
  • What is the max cluster size? The upstream-tested envelope is 5,000 nodes, 150,000 pods, and 10,000 Services. See infrastructure/kubernetes/index#Scale Limits (Upstream) and Reference.
  • Is Docker still supported? Dockershim was removed in v1.24. containerd and CRI-O are the supported runtimes. Images built with Docker (OCI) still run.
  • What happened to ingress-nginx? It was retired in March 2026. Migrate to Gateway API controllers (see How-to Guides).
  • How does scheduling work? Filter, then Score, then Bind, through scheduling-framework plugins. See infrastructure/kubernetes/explanation#How It Works.
  • Which Kubernetes versions are supported right now (2026-09)? 1.37, 1.36, and 1.35. 1.34 is in maintenance mode until EOL on 2026-10-27. See Reference.