Explanation¶
OpenStack is a distributed cloud operating system composed of independent services. Each service exposes its own REST API, keeps its own database schema, and talks to its own worker processes over a message bus (RabbitMQ via oslo.messaging). Keystone ties them together with a shared identity and service catalog. This service-oriented design allows incremental adoption and huge scale, but costs far more operational effort than single-daemon platforms like OpenNebula or Proxmox VE.
This page explains how the pieces work and why. Look-up facts (release matrix, ports, drivers, retirements, hardening checklist) live in Reference; procedures live in How-to Guides; the hub is OpenStack.
Component Overview¶
The diagram below shows the main control-plane services, the shared infrastructure they depend on, and the agents on compute nodes. Solid arrows are REST calls, dotted arrows are RPC over RabbitMQ.
flowchart TB
subgraph Clients["Clients"]
HORIZON["Horizon / Skyline<br/>(web UI)"]
OSC["openstack CLI<br/>(openstacksdk)"]
end
subgraph API["Control plane (API services behind HAProxy)"]
KEYSTONE["Keystone<br/>(identity, catalog)"]
NOVAAPI["nova-api"]
PLACEMENT["Placement"]
NEUTRON["neutron-server<br/>(ML2/OVN)"]
CINDER["cinder-api / cinder-scheduler"]
GLANCE["glance-api"]
HEAT["heat-api / heat-engine"]
end
subgraph NovaCtl["Nova control"]
CONDUCTOR["nova-conductor"]
SCHED["nova-scheduler"]
end
subgraph Shared["Shared infrastructure"]
MQ["RabbitMQ"]
DB["MariaDB / Galera"]
MC["Memcached"]
OVNDB["OVN NB / SB databases"]
end
subgraph Hosts["Compute and storage hosts"]
COMPUTE["nova-compute<br/>(libvirt, KVM)"]
OVNCTL["ovn-controller<br/>(Open vSwitch)"]
CVOL["cinder-volume"]
CEPH["Ceph RBD"]
end
HORIZON --> KEYSTONE
OSC --> KEYSTONE
OSC --> NOVAAPI
NOVAAPI --> PLACEMENT
SCHED --> PLACEMENT
NOVAAPI -.-> CONDUCTOR
CONDUCTOR -.-> SCHED
CONDUCTOR -.-> COMPUTE
NOVAAPI --- MQ
COMPUTE --> NEUTRON
COMPUTE --> GLANCE
COMPUTE --> CINDER
NEUTRON --> OVNDB
OVNDB --> OVNCTL
CINDER -.-> CVOL
CVOL --> CEPH
COMPUTE --> CEPH
HEAT --> NOVAAPI
NOVAAPI --> DB
NEUTRON --> DB
CINDER --> DB
KEYSTONE --> DB
KEYSTONE --> MC
Core Services¶
Keystone (Identity)¶
Keystone is the central authentication, authorization and service catalog for OpenStack:
- Authentication: passwords, tokens, application credentials, TOTP, and federated identity (SAML 2.0, OpenID Connect) mapped onto local groups and roles
- Authorization: role assignments on projects, domains or the whole system; each service enforces them through its own policy rules (oslo.policy)
- Service Catalog: registry of every service endpoint (public, internal, admin interfaces per region)
- Multi-domain: separate identity backends per domain (SQL, LDAP / Active Directory)
- Token types: Fernet (default, symmetric-key encrypted) or JWS (asymmetric signatures); neither is stored in the database
- Every API request carries a token that the receiving service validates (with caching through keystonemiddleware and Memcached)
Nova (Compute)¶
Nova manages the lifecycle of virtual machine instances:
- nova-api: accepts and validates REST requests (microversioned API)
- nova-scheduler: asks Placement for candidate hosts, then filters and weighs them
- nova-conductor: orchestrates build, resize and migration tasks and mediates database access for compute nodes (compute hosts never touch the database directly)
- nova-compute: runs on each hypervisor node and drives libvirt/KVM (or the Ironic, VMware or z/VM drivers)
- nova-novncproxy / nova-spicehtml5proxy / nova-serialproxy: console access
- Cells v2: large clouds shard instances into cells, each with its own database and message queue, under a shared API database
- Supports live migration, resize, snapshot, shelve and evacuate
Placement¶
Placement started inside Nova and became a separate project. It tracks resource providers (compute nodes, shared storage pools, bandwidth on physical NICs), their inventories (VCPU, MEMORY_MB, DISK_GB, custom resource classes) and traits. The scheduler asks Placement for allocation candidates before it filters hosts, and claims resources atomically, which removes most scheduling races.
Neutron (Networking)¶
Neutron provides network connectivity as a service:
- Networks and subnets: virtual L2 segments (VLAN, VXLAN, Geneve) with IP address management
- Routers: virtual L3 routing between subnets, with SNAT for external access
- Floating IPs: one-to-one NAT from external networks to ports
- Security groups: per-port stateful firewall rules
- ML2 plugin: mechanism drivers for OVN (the DevStack default), Open vSwitch, SR-IOV and vendor SDNs. The Linux bridge driver was removed in 2025.1 Epoxy
- With ML2/OVS: L2 agent per host plus L3, DHCP and metadata agents
- With ML2/OVN: neutron-server writes to the OVN Northbound database;
ovn-northdcompiles logical flows into the Southbound database;ovn-controlleron each chassis programs Open vSwitch. Routing, DHCP and security groups are distributed, and the OVN agent replaced the separate metadata agent in 2025.2
Cinder (Block Storage)¶
Cinder provides persistent block storage volumes:
- Volume creation, attachment, snapshot, cloning, backup, retype and migration
- Backend drivers: LVM, Ceph RBD, NFS and dozens of vendor arrays
- Volume types with extra specs (performance tiers, replication, encryption)
- Availability zone awareness for volume placement
- Backups to Swift, Ceph, NFS or POSIX targets
Glance (Image)¶
Glance manages VM disk images:
- Stores, discovers and serves bootable images
- Disk formats: QCOW2, RAW, VHD, VHDX, VMDK, VDI, ISO, PLOOP, AKI/ARI/AMI
- Backend stores: file, Ceph RBD, Swift, S3, Cinder, HTTP (read-only); multiple stores can be enabled at once
- Image properties drive scheduling and hardware features (for example
hw_firmware_type,hw_mem_encryption_model) - Image sharing across projects, and interoperable image import with conversion
Swift (Object Storage)¶
Swift provides highly available, eventually consistent object storage:
- Accounts, containers and objects; native Swift API plus an S3-compatible API through the
s3apimiddleware - Ring-based data placement across zones and regions, replication or erasure coding
- Large objects (segmented uploads), versioning and expiration
- Account and container listings live in per-partition SQLite databases replicated by Swift itself, not in the shared MariaDB cluster
Orchestration and Supporting Services¶
Heat (Orchestration)¶
Heat provisions OpenStack resources from declarative templates (HOT, the Heat Orchestration Template format, plus a CloudFormation-compatible API):
- Define stacks of interconnected resources (servers, networks, volumes, load balancers)
- Autoscaling groups driven by Aodh alarms (Telemetry)
- Dependency ordering, updates in place and rollback on failure
- Environment files separate parameter values from templates
Horizon and Skyline (Dashboards)¶
Horizon is the long-standing Django web UI, extensible through plugins (for example heat-dashboard, magnum-ui). Skyline is a newer official dashboard (a Python API server plus a single-page console) aimed at a faster, more modern user experience.
Magnum (Container Infrastructure)¶
Magnum provisions Kubernetes clusters on OpenStack. The in-tree driver (k8s_fedora_coreos_v1) builds clusters with Heat templates; the Swarm and Mesos drivers are gone. The Magnum team now also maintains magnum-capi-helm, a driver that uses Cluster API (via Helm charts) to manage cluster lifecycle, and a vendor alternative (magnum-cluster-api) exists as well. Many operators instead run Cluster API with the OpenStack provider (CAPO) directly.
Ironic (Bare Metal)¶
Ironic provisions physical servers through PXE/iPXE or Redfish virtual media, with in-band (agent) or out-of-band (Redfish) inspection. It works behind Nova (the ironic.IronicDriver makes a bare metal node look like a flavor-sized host) or standalone. Since 2026.1 a standalone networking service can configure switch ports for bare metal without Neutron.
Barbican (Key Manager)¶
Barbican stores secrets, symmetric and asymmetric keys, and certificates. Cinder and Nova use it (through the castellan library) for LUKS encryption keys, Octavia for TLS certificates, and Swift optionally for its encryption root secret. Backends are pluggable: the software simple_crypto plugin, PKCS#11 HSMs, KMIP devices, HashiCorp Vault and Dogtag. See the backend table.
Shared Infrastructure¶
RabbitMQ (Message Bus)¶
Services use RabbitMQ through oslo.messaging for RPC between their own components (API to conductor, conductor to compute) and for notifications consumed by other services:
- AMQP 0-9-1, topic exchanges per service (for example
nova,notifications.info) - RPC is either a
cast(fire and forget) or acall(waits for a reply queue) - Classic mirrored queues were removed in RabbitMQ 4.0; current deployment tools move to quorum queues (and streams for fanout)
- Each service normally uses its own vhost or user
MariaDB / Galera (Database)¶
Each service owns one or more schemas (for example keystone, nova_api, nova_cell0, nova, placement, neutron, cinder, glance, heat), typically in one MariaDB Galera cluster:
- Galera provides synchronous multi-primary replication
- HAProxy usually sends all writes to one node at a time to avoid certification conflicts
- Schema changes happen through each project's
*-manage db sync(alembic migrations)
Memcached¶
Memcached caches Keystone tokens, validation results and catalog lookups (keystonemiddleware memcached_servers, oslo.cache backends), which removes most repeat load on Keystone and the database.
Request Flow: Instance Creation¶
The sequence below shows openstack server create on a KVM cloud with ML2/OVN and a Glance image. RPC hops go through RabbitMQ; REST hops carry the user's token (or a service token).
sequenceDiagram
actor User
participant CLI as openstack CLI
participant KS as Keystone
participant API as nova-api
participant Cond as nova-conductor
participant Sched as nova-scheduler
participant PL as Placement
participant Cmp as nova-compute
participant GL as Glance
participant NE as Neutron
participant Lv as libvirt / QEMU
User->>CLI: server create --flavor --image --network
CLI->>KS: POST /v3/auth/tokens
KS-->>CLI: token + service catalog
CLI->>API: POST /servers
API->>KS: validate token (cached)
API->>API: check quota, create BuildRequest
API--)Cond: RPC cast schedule_and_build_instances
API-->>CLI: 202 Accepted, status BUILD
Cond->>Sched: RPC call select_destinations
Sched->>PL: GET allocation_candidates
PL-->>Sched: candidate hosts
Sched->>Sched: filter and weigh hosts
Sched->>PL: claim allocations
Sched-->>Cond: selected host
Cond--)Cmp: RPC cast build_and_run_instance
Cmp->>GL: download image (or RBD clone)
Cmp->>NE: create or bind port
NE-->>Cmp: port details (MAC, IP)
Cmp->>Lv: define and start domain
NE--)Cmp: network-vif-plugged event
Cmp->>Cond: update instance ACTIVE
Release Model and SLURP¶
OpenStack releases every six months. Until 2022, upgrades were only tested between adjacent releases, so a one-year upgrade cadence meant a "fast-forward upgrade" through an intermediate release. The TC's 2022 release-cadence resolution introduced SLURP (Skip Level Upgrade Release Process): every .1 release (2023.1 Antelope, 2024.1 Caracal, 2025.1 Epoxy, 2026.1 Gazpacho, 2027.1 Indri) is a SLURP, and upgrades from one SLURP directly to the next are tested (grenade skip-level jobs) and supported. Deprecation, waiting and removal may only land in SLURP releases, so an operator jumping SLURP to SLURP never misses a required change that was introduced in the skipped release.
A release moves through a fixed lifecycle; the dates for each series are in the release matrix.
stateDiagram-v2
[*] --> Development: cycle opens on master
Development --> Maintained: coordinated release (stable branch)
Maintained --> Unmaintained: SLURP release, about 18 months after GA
Maintained --> EndOfLife: non-SLURP release, about 18 months after GA
Unmaintained --> EndOfLife: no volunteers keep CI green
EndOfLife --> [*]: branch deleted, series-eol tag
Trade-off: operators on a yearly cadence get a supported jump, but they still run each project's data migrations for the skipped release, and the surrounding dependencies (RabbitMQ, MariaDB, Ceph, the host OS) do not follow SLURP. Kolla-Ansible, for example, requires an intermediate RabbitMQ upgrade before a skip-level upgrade.
Governance and Community¶
- 2010: Rackspace and NASA launched OpenStack (Swift and Nova origins).
- 2012: the independent OpenStack Foundation took over governance.
- 2020: the foundation renamed itself the Open Infrastructure (OpenInfra) Foundation and began hosting other projects (Kata Containers, StarlingX, Zuul).
- 2025: the OpenInfra Foundation announced on 2025-03-12 its intent to join the Linux Foundation, and completed the move in mid-2025. OpenStack's technical governance (the Technical Committee, project teams and PTLs, elections, the four opens) is unchanged; the Linux Foundation now provides the legal, financial and event umbrella.
Technical work happens on OpenDev (Gerrit code review, Zuul CI); the GitHub openstack/* repositories are read-only mirrors.
Eventlet Removal¶
Most OpenStack services historically used eventlet green threads with monkey-patching for concurrency. Eventlet has few maintainers and breaks with new CPython versions, so the TC selected a community goal to remove it. Progress so far:
- Neutron (2025.2 Flamingo): API, RPC workers and agents run in native threading mode.
- Nova (2026.1 Gazpacho): nova-api, nova-scheduler and nova-metadata use native threads by default; nova-conductor and nova-compute can opt in via
OS_NOVA_DISABLE_EVENTLET_PATCHING=true. Nova's own docs still call native threading experimental. - Other projects (for example Watcher, Designate, Cyborg, Manila) have native-threading modes in varying stages.
For operators the practical effect is new thread-pool tunables (for example Nova's executor_thread_pool_size and cell_worker_thread_pool_size) and a need to re-check memory sizing during upgrades.
Deployment Models¶
OpenStack itself only ships Python packages; deployment tools decide how the services run:
| Model | Examples | Why teams choose it |
|---|---|---|
| Containers on hosts, driven by Ansible | Kolla-Ansible (+ Kayobe) | Immutable images, simple rollbacks, no Kubernetes dependency |
| Ansible roles on hosts or LXC | OpenStack-Ansible | Source-based installs, fine-grained control |
| Control plane on Kubernetes | OpenStack-Helm, Genestack, Canonical Sunbeam, Red Hat RHOSO | Reuse Kubernetes for HA, rolling updates and operators |
| Development | DevStack | Single-node developer environment only |
TripleO (OpenStack deployed by OpenStack) was retired in January 2024; Red Hat moved its product to OpenShift operators (RHOSO).
Neutron Networking (OVN)¶
With ML2/OVN, tenant traffic stays in Open vSwitch flows on each chassis; there are no network nodes running L3 agents. Gateway chassis only handle north-south SNAT and floating IPs for centralized cases (distributed floating IPs are optional).
flowchart TB
subgraph Tenant["Tenant network (Geneve overlay)"]
VM1["VM 1<br/>(10.0.0.2)"]
VM2["VM 2<br/>(10.0.0.3)"]
end
subgraph OVN["OVN logical topology"]
LS["Logical switch<br/>(tenant subnet)"]
LR["Logical router<br/>(distributed)"]
GW["Gateway chassis port<br/>(SNAT, floating IPs)"]
end
subgraph Physical["Provider network"]
ExtNet["br-ex / provider bridge<br/>(external VLAN or flat)"]
end
VM1 --> LS
VM2 --> LS
LS --> LR
LR --> GW
GW --> ExtNet
Ceph Integration¶
Ceph is the most common storage backend. Using RBD for Glance, Cinder and Nova ephemeral disks lets new instances boot from copy-on-write clones instead of downloading images, and makes live migration a memory-only operation.
| OpenStack Service | Ceph Layer | Purpose |
|---|---|---|
| Cinder | RBD (RADOS Block Device) | Persistent volumes |
| Glance | RBD | Image storage; enables copy-on-write clones |
| Nova | RBD (images_type = rbd) |
Ephemeral disks on shared storage |
| Swift API | RGW (RADOS Gateway) | Swift and S3-compatible object API without running Swift |
| Manila | CephFS | Shared file systems |
Ceph is versioned independently: as of 2026-09, Tentacle (20.2.x) and Squid (19.2.x, target EOL 2026-10-31) are the active Ceph releases.
Multi-Region Architecture¶
Regions are complete OpenStack deployments that share a Keystone (and usually Horizon), so one set of credentials sees every region's endpoints in the catalog. Within a region, Nova cells and availability zones shard scale and failure domains.
flowchart TB
subgraph Global["Shared services"]
KS_G["Keystone<br/>(shared identity, catalog with region endpoints)"]
end
subgraph Region1["RegionOne"]
Nova1["Nova API + cells"]
Neutron1["Neutron"]
Cinder1["Cinder"]
Compute1["Compute hosts"]
end
subgraph Region2["RegionTwo"]
Nova2["Nova API + cells"]
Neutron2["Neutron"]
Cinder2["Cinder"]
Compute2["Compute hosts"]
end
KS_G --> Nova1
KS_G --> Nova2
Nova1 --> Compute1
Nova2 --> Compute2
Security Model¶
Authentication: Keystone¶
Every user, service and application authenticates to Keystone and receives a scoped token. Services validate tokens through keystonemiddleware, then evaluate the request against their policy rules.
flowchart TB
subgraph Callers["Callers"]
Admin["Cloud admin<br/>(system scope)"]
Tenant["Project member<br/>(project scope)"]
App["Application credential"]
end
subgraph Keystone["Keystone"]
AuthN["Auth plugins<br/>(password, TOTP, OIDC, SAML)"]
Token["Token provider<br/>(Fernet or JWS)"]
Catalog["Service catalog"]
end
subgraph Services["Service APIs"]
MW["keystonemiddleware<br/>(auth_token)"]
Policy["oslo.policy rules<br/>(policy.yaml overrides)"]
Nova["Nova / Neutron / Cinder / Glance"]
end
Admin --> AuthN
Tenant --> AuthN
App --> AuthN
AuthN --> Token
Token --> Catalog
Catalog --> MW
MW --> Policy
Policy --> Nova
Authorization: RBAC and Policy¶
OpenStack uses a policy engine (oslo.policy, with optional policy.yaml overrides per service) to enforce RBAC. The "secure RBAC" (SRBAC) community goal introduced scoped default roles:
| Default Role | Scope | Typical Permissions |
|---|---|---|
admin |
System or project | Full administrative access |
manager |
Project | Project-level management tasks without full admin (for example some migration operations in Nova since 2025.2) |
member |
Project | Create and manage resources in own project |
reader |
System or project | Read-only access |
service |
Project | Service-to-service API calls |
Project Isolation¶
Projects provide the multi-tenancy boundary:
- Each project has isolated resources (instances, networks, volumes, private images)
- Users can hold different roles in different projects
- Quotas cap resources per project
- Tenant networks are isolated by VLAN IDs or VXLAN/Geneve VNIs, and by per-project routers
Network Security: Neutron¶
Security groups are distributed stateful firewalls applied per port (in OVN as ACLs, in ML2/OVS via the OVS firewall driver):
- Each project gets a
defaultgroup that allows all egress and allows ingress only from members of the same group; anything else is denied - Connection tracking allows return traffic for established flows (stateless groups are possible)
- Rules can reference remote groups instead of CIDRs
- Port security (anti-spoofing of MAC/IP) is on by default and can be disabled per port or network, for example for virtual appliances
| Mechanism | Level | Use Case |
|---|---|---|
| Projects | L2/L3 | Tenant isolation |
| VLANs | L2 | Simple segmentation, provider networks |
| VXLAN/Geneve | L2 overlay | Scalable multi-tenant isolation |
| SR-IOV | L2 | Hardware NIC partitioning (bypasses OVS security groups) |
| OVS-DPDK | L2/L3 | High-performance userspace switching |
Encryption¶
TLS should protect every hop: public and internal API endpoints (usually terminated on HAProxy), RabbitMQ connections, database connections and Galera replication, and Memcached where supported. Data at rest uses LUKS (Cinder volumes, Nova ephemeral disks) with keys held in Barbican, and Swift's encryption middleware; the per-service matrix is in Reference.
Audit¶
Keystone and the other services can emit CADF (Cloud Auditing Data Federation) events through the audit middleware, which records who did what to which resource. Forward these and service logs to a central store (Loki, OpenSearch) for retention and alerting. The actionable hardening checklist and known pitfalls tables are in Reference.
Sources¶
- Nova architecture
- Nova service concurrency (eventlet removal)
- Placement documentation
- Neutron OVN architecture
- Neutron eventlet deprecation reference
- Cinder Ceph RBD driver
- Keystone token providers
- TC goal: Remove Eventlet
- TC resolution: Release Cadence Adjustment (SLURP)
- Linux Foundation press release: OpenInfra intent to join (2025-03-12)
- OpenInfra blog: We've Closed the Deal
- Ceph releases
- OpenStack Security Guide