Infrastructure¶
Domain summary
The compute foundation under everything else in this knowledge base: container engines and orchestrators (Docker Engine 29.8.1, Kubernetes v1.37.1), private virtualization and cloud platforms (Proxmox VE 9.2, OpenNebula 7.4.1, OpenStack 2026.1 Gazpacho), the four public clouds covered in depth (AWS, Google Cloud, Alibaba Cloud, Tencent Cloud), and two cross-cutting disciplines: multi-cloud governance (landing zones, policy as code, FinOps) and AI platform engineering (GPUs on Kubernetes with DRA, gang schedulers and LLM serving). The platforms sit at different layers and are usually combined, not chosen as rivals. All facts on this page come from the topic pages, refreshed on 2026-09-25.
Topics¶
| Topic | What it is | Current version (checked 2026-09-25) | License |
|---|---|---|---|
| Kubernetes | Container orchestrator: declarative desired state reconciled by controllers, with scheduling, service discovery, storage and self-healing. First CNCF graduated project. DRA, in-place pod resize and native sidecars are GA; Ingress NGINX is retired in favour of Gateway API | v1.37.1 (v1.37.0 "Garhwal" 2026-08-26); 1.37, 1.36 and 1.35 supported | Apache-2.0 |
| Docker | OCI container toolchain: the open-source Engine (Moby, dockerd on containerd and runc), BuildKit/Buildx, Compose and Swarm mode, plus the proprietary Docker Desktop, Docker Hub and Hardened Images |
Engine 29.8.1 (2026-09-15); Desktop 4.92.0; Compose v5.5.1 | Engine, CLI, Compose, BuildKit: Apache-2.0. Desktop: proprietary, paid for companies with more than 250 employees or more than $10M revenue |
| AI Platform Engineering | GPU infrastructure for training and LLM inference on Kubernetes: GPU Operator and DRA drivers, Kueue, Volcano and KAI Scheduler for gang scheduling and quotas, Ray, and serving with vLLM, KServe, llm-d and NVIDIA Dynamo behind the Gateway API Inference Extension | Discipline. Baseline K8s 1.37; GPU Operator v26.7.1, KAI Scheduler v0.18.0, vLLM 0.30.0, Dynamo 1.5.0 | Core projects Apache-2.0 (Run:ai is the main commercial layer) |
| Proxmox VE | Debian-based hypervisor appliance: QEMU/KVM VMs and LXC containers, corosync cluster with HA and the CRS load balancer, ZFS and Ceph, SDN, and integrated Proxmox Backup Server. Datacenter Manager 1.1 adds a multi-cluster view | 9.2 (2026-05-21; arm64 2026-08-05). PVE 8 reached EOL in 2026-08 | AGPLv3; optional support subscriptions EUR 120-1,100 per socket-year |
| OpenNebula | Cloud manager with one central daemon (oned) over KVM and LXC hosts: self-service portal, groups and VDCs, federated zones, OneForm edge-cluster provisioning, OneDRS (ILP-based) load balancing, OneKS Kubernetes and NVIDIA AI-factory integrations |
7.4.1 (2026-09-10); 7.4 "Helix" released 2026-07-27 | Community Edition Apache-2.0; Enterprise Edition under commercial subscription terms |
| OpenStack | Cloud operating system for multi-tenant IaaS: cooperating Python services (Nova, Neutron, Cinder, Glance, Keystone, Ironic and more) glued by RabbitMQ and MariaDB/Galera. Governed by its Technical Committee under the OpenInfra Foundation, part of the Linux Foundation since 2025 | 2026.1 "Gazpacho" (2026-04-01, SLURP); 2026.2 "Hibiscus" planned for 2026-09-30 | Apache-2.0 |
| AWS | Amazon's public cloud, from one VPC to enterprise landing zones: Organizations with SCPs, RCPs and declarative policies, Control Tower, IAM Identity Center, Transit Gateway, Cloud WAN, PrivateLink, VPC Lattice and multi-Region DR | Not versioned. 39 Regions / 123 AZs (2026-06); Control Tower landing zone 4.0; LZA 1.15.5 | Proprietary service; CLI, CDK and SDKs Apache-2.0 |
| GCP | Google's public cloud: resource hierarchy (organization, folders, projects), organization policies, Shared VPC, Network Connectivity Center, Private Service Connect, Cloud NGFW, VPC Service Controls, and the Fabric FAST and enterprise foundations blueprints | Not versioned. 43 regions / 130 zones (2026-09); Fabric FAST v58.0.0; terraform-example-foundation v6.0.0 | Proprietary service; Fabric and CFT modules Apache-2.0, Terraform provider MPL-2.0 |
| Alibaba Cloud | Largest cloud in mainland China with fast growth in Southeast Asia: Resource Directory with control policies and the Agentic Cloud Governance Center landing zone, CEN with Transit Routers, PolarDB and PolarDB-X, and the Qwen AI stack | Not versioned. 31 regions / 107 AZs (2026-09-23); Terraform aliyun/alicloud 1.293.0 |
Proprietary service; CLI Apache-2.0, Terraform provider MPL-2.0 |
| Tencent Cloud | Third-largest cloud in mainland China: Tencent Cloud Organization (TCO) with SCPs, Identity Center and the Control Center landing zone, CCN transit networking, the TDSQL database family, and Anti-DDoS and EdgeOne edge security | Not versioned. 23 regions / 66 AZs (2026-08-18); Terraform tencentcloud 1.83.33 |
Proprietary service; CLI and SDKs Apache-2.0, Terraform provider MPL-2.0 |
| Multi-Cloud Governance | Operating model for estates on two or more clouds: one intent in Git, enforced in CI, at Terraform runs, at each cloud's organization guardrails and at Kubernetes admission; plus identity federation, networking, landing zones, OpenTelemetry and FOCUS-based FinOps | Pattern. FOCUS 1.4; OPA 1.21.0; Kyverno 1.19.1; Crossplane 2.4.2 | Tools mostly Apache-2.0; Terraform BSL 1.1, OpenTofu MPL-2.0 |
How the Domain Fits Together¶
The map places each topic on its layer: private platforms and public clouds provide machines, Kubernetes schedules containers built with Docker on top of either, AI platform engineering extends Kubernetes for GPUs, and multi-cloud governance spans the public clouds.
flowchart TB
subgraph PRIV["Private infrastructure (KVM hosts)"]
PVE["Proxmox VE 9.2<br/>hypervisor appliance"]
ONE["OpenNebula 7.4<br/>cloud manager"]
OS["OpenStack 2026.1<br/>cloud OS"]
end
subgraph PUB["Public clouds"]
AWS["AWS<br/>Organizations, Control Tower"]
GCP["GCP<br/>org, folders, FAST"]
ALI["Alibaba Cloud<br/>Resource Directory"]
TC["Tencent Cloud<br/>TCO, Control Center"]
end
MCG["Multi-Cloud Governance<br/>policy as code, FOCUS, landing zones"]
K8S["Kubernetes v1.37<br/>EKS, GKE, ACK, TKE, Magnum, OneKS"]
DOCKER["Docker Engine 29.8<br/>builds OCI images"]
AIPE["AI Platform Engineering<br/>DRA, Kueue, KAI, vLLM, llm-d"]
PRIV -->|"VMs and bare metal"| K8S
PUB -->|"managed control planes"| K8S
DOCKER -->|"OCI images run on containerd"| K8S
K8S --> AIPE
MCG -.->|"guardrails, identity, cost"| PUB
Comparisons¶
| Comparison | Scope |
|---|---|
| Infrastructure Platforms Comparison | Docker vs Kubernetes vs OpenStack vs OpenNebula across stack layers: architecture, scale, operations, security, cost and use-case fit, with a decision flowchart |
| Proxmox VE vs OpenNebula vs OpenStack | The three private-infrastructure platforms on an ambition axis: architecture, operations, tenancy, backup, cost and weighted decision matrices for VMware replacement and multi-tenant cloud |
| Public Cloud Landing Zones | AWS vs Google Cloud vs Alibaba Cloud vs Tencent Cloud: account hierarchy, guardrails, landing-zone services, identity, transit networking, audit and IaC, with a service mapping and a decision flowchart |
All comparison notes are listed in the comparisons index.
When to Use Which¶
These rules match the decision flowcharts in the comparison pages.
| Need | Reach for | Details |
|---|---|---|
| Build and run containers on one host, local development, CI image builds | Docker (Compose for multi-container apps) | Platforms comparison |
| Simple multi-host container clusters without Kubernetes | Docker Swarm mode (maintained, but niche) | Docker |
| Run containerized services at scale with self-healing and autoscaling | Kubernetes (managed EKS, GKE, ACK or TKE where possible) | Kubernetes |
| GPU training and LLM inference on shared clusters | Kubernetes plus the AI platform stack (DRA, Kueue or KAI, vLLM or llm-d) | AI Platform Engineering |
| Replace VMware on a few to tens of nodes with one infra team | Proxmox VE | Proxmox vs OpenNebula vs OpenStack |
| Cloud-style self-service, VDC tenancy, edge sites or GPU factories without a large platform team | OpenNebula | Proxmox vs OpenNebula vs OpenStack |
| Multi-tenant IaaS at hundreds to thousands of nodes, telco NFV, bare metal as a service | OpenStack (often through a vendor distribution) | OpenStack |
| Global public-cloud footprint with the broadest service catalogue | AWS or Google Cloud | Landing zones comparison |
| Workloads in mainland China, or a China plus Southeast Asia footprint | Alibaba Cloud or Tencent Cloud (separate China-site accounts, ICP filing) | Landing zones comparison |
| Two or more clouds under one set of guardrails and cost reporting | Multi-cloud governance patterns | Multi-Cloud Governance |
Landscape¶
The landscape is converging on containers as the main compute primitive while still running VMs, bare metal and serverless workloads. Kubernetes is the de facto orchestration standard, and most teams avoid running control planes themselves: managed services (EKS, GKE, AKS, ACK, TKE) and lightweight distributions (k3s, Talos) cover most clusters. In 2025-2026 the project made Dynamic Resource Allocation GA (1.34) and retired Ingress NGINX (March 2026) in favour of Gateway API, which matters both for AI clusters and for every ingress migration.
The VMware licensing changes after Broadcom's acquisition pushed thousands of vSphere estates to re-platform in 2024-2026. Proxmox VE absorbed much of the small and mid-size demand; OpenNebula and OpenStack compete where self-service, multi-tenancy or scale is needed. OpenStack's foundation, the OpenInfra Foundation, joined the Linux Foundation in 2025, while the project keeps retiring peripheral services and concentrating on the core.
Sovereignty is the other driver. EU, Chinese and Southeast Asian data-residency rules favour on-premises OpenStack, OpenNebula and Proxmox deployments, and the hyperscalers answer with partitions such as the AWS European Sovereign Cloud (GA 2026-01-15). Alibaba Cloud and Tencent Cloud keep adding regions in Southeast Asia and the Middle East, which makes China-plus-ASEAN multi-cloud estates common.
Container-VM convergence
Kata Containers, Firecracker and KubeVirt give VM-level isolation with container-like workflows. KubeVirt runs traditional VMs next to pods in the same Kubernetes cluster, so VM estates can migrate gradually instead of in one cut-over.
Multi-cloud estates abstract provider primitives with tools such as Crossplane, Cluster API and internal platform layers. Cluster API provides a declarative Kubernetes API for creating and upgrading clusters, with providers for the major clouds, vSphere, OpenStack and bare metal.
Key Concepts¶
Container Orchestration¶
The automated management of container lifecycle, scheduling, networking, storage and self-healing across a cluster of nodes. Kubernetes implements this through a declarative desired-state model: users submit manifests, and controllers continuously reconcile actual state to match.
Key scheduling decisions include:
- Resource requests and limits: guarantee minimum CPU and memory (requests) and cap consumption (limits)
- Affinity and anti-affinity: co-locate or spread pods based on node labels or existing pod placement
- Topology spread constraints: distribute pods evenly across zones, regions or custom topology domains
- Priority and preemption: higher-priority pods evict lower-priority pods when resources are scarce
- Device allocation (DRA) and gang scheduling: GPUs and other devices are claimed through ResourceClaims, and all-or-nothing placement for distributed jobs comes from Kueue, Volcano, KAI Scheduler or the Workload API (beta in 1.37)
Infrastructure Abstraction¶
Decoupling application teams from infrastructure details through APIs, CLIs and internal developer platforms. Crossplane extends Kubernetes with Composite Resource Definitions (XRDs) that expose simplified infrastructure APIs, and Cluster API manages Kubernetes clusters themselves as resources. The goal is self-service: developers request a database, a cluster or a network without knowing the provider specifics.
Landing Zones¶
Landing zone components
A landing zone is a pre-configured multi-account environment that implements an organization's governance, networking, identity and security baseline before workloads arrive. Every cloud offers the same building blocks under different names: an account hierarchy (AWS Organizations OUs, Google Cloud folders and projects, Alibaba Resource Directory folders, Tencent TCO departments), preventive guardrails (SCPs and RCPs, Organization Policy, control policies, TCO SCPs), and a packaged setup service or blueprint (AWS Control Tower, Fabric FAST or the enterprise foundations blueprint, Agentic Cloud Governance Center, Control Center). Azure uses management groups, Azure Policy and Azure Landing Zones. The Public Cloud Landing Zones comparison maps them side by side.
Network Fabric¶
The interconnection layer that provides reachability between workloads across nodes, zones, regions and clouds. In Kubernetes, CNI plugins (Cilium, Calico, Flannel) provide pod-to-pod connectivity. In the cloud, hub-and-spoke transit services connect VPCs: AWS Transit Gateway and Cloud WAN, Google Cloud Network Connectivity Center, Alibaba CEN with Transit Routers, and Tencent CCN. Dedicated lines (Direct Connect, Cloud Interconnect, Express Connect, Tencent Direct Connect) and SD-WAN bridge on-premises and cloud networks. Overlays (VXLAN, Geneve, WireGuard) encapsulate traffic across underlays.
Control Plane vs Data Plane¶
A basic pattern in distributed systems: decision-making components (control plane) are separated from traffic-handling components (data plane). In Kubernetes, the control plane is the API server, scheduler, controller manager and etcd; the data plane is the kubelet, kube-proxy and container runtime on each worker node. The separation allows independent scaling, security isolation and separate failure domains. Service meshes (Istio's istiod vs Envoy sidecars or ambient ztunnel), Gateway API implementations (for example the Envoy Gateway controller vs its Envoy proxies) and cloud networking (SDN controllers vs forwarding elements) follow the same pattern.
Related Domains¶
- Networking: CNI plugins for Kubernetes, see the CNI comparison
- Service Mesh: Istio, Linkerd and Envoy Gateway (Gateway API)
- Storage: Ceph (used by Proxmox VE, OpenStack and OpenNebula), Longhorn and object storage
- IaC: Terraform, OpenTofu and Pulumi provision the clouds and platforms in this domain
- CI/CD: Argo CD and Flux deploy onto Kubernetes clusters
- Observability: OpenTelemetry and metrics, logs and traces backends for clusters and clouds
- Secrets: Vault, External Secrets Operator and SOPS
- LLM Inference: model-serving techniques that run on the AI platform stack
Sources¶
- Kubernetes and Kubernetes releases
- Docker documentation and Docker Engine release notes
- Proxmox VE and PVE roadmap
- OpenNebula and OpenNebula documentation
- OpenStack releases and OpenInfra Foundation joins the Linux Foundation
- AWS global infrastructure and AWS Control Tower
- Google Cloud locations and Google Cloud landing zone design
- Alibaba Cloud and Agentic Cloud Governance Center
- Tencent Cloud and Tencent Cloud Control Center
- FinOps FOCUS specification
- Per-topic sources are listed on each topic page.
Open Questions¶
- As Kubernetes dominates orchestration, what role will OpenStack play in five years? Will it retreat to telco and sovereign niches, or will Ironic (bare metal) and Kubernetes-hosted deployments (OpenStack-Helm, Genestack, Sunbeam) keep it relevant as a substrate?
- How do organizations implement multi-cloud without a lowest-common-denominator abstraction that gives up each provider's distinctive capabilities?
- Will WebAssembly runtimes such as SpinKube be scheduled natively by Kubernetes next to containers, or will a separate orchestration layer emerge?
- Did OpenStack 2026.2 "Hibiscus" ship on its planned date (2026-09-30)? Update the version rows here and in the comparisons when it does.