Explanation¶
Scope
How OpenNebula works and why it is built that way: the centralized oned control plane, the driver framework, the 7.x scheduler and OneDRS, the service layer (Sunstone, OneFlow, OneGate, OneForm, OneKS), storage and networking models, the security model, performance characteristics, and the European sovereign-cloud context behind the 7.x roadmap. Look-up tables (ports, drivers, config keys, VM states) live in Reference; commands live in How-to Guides.
OpenNebula is a centralized cloud management platform. Its core runs as one C++ daemon (oned) on a front-end node, which keeps all state in a SQL database and drives hypervisor hosts, storage and networks through pluggable driver scripts executed over SSH. This is the opposite of OpenStack's design of many cooperating services, and it is the main reason OpenNebula is considered simpler to install and operate.
See also: OpenNebula hub, Reference, How-to Guides.
Architecture¶
Component Overview¶
The diagram shows the 7.4 front-end services, the driver framework, and what each driver talks to on the hosts.
graph TD
subgraph Frontend["Front-end node"]
ONED["oned<br/>(core daemon, XML-RPC :2633, gRPC :2634)"]
DB[("SQLite or MariaDB/MySQL")]
SCHED["one_sched driver<br/>(rank + OneDRS)"]
MON["Monitor daemon<br/>(:4124)"]
HEM["Hook Execution Manager"]
FIREEDGE["FireEdge / Sunstone<br/>(:2616)"]
ONEFLOW["OneFlow (:2474)"]
ONEGATE["OneGate (:5030)"]
ONEFORM["OneForm<br/>(cluster provisioning)"]
ONEKS["OneKS<br/>(Kubernetes as a Service)"]
end
subgraph Drivers["Driver framework (/var/lib/one/remotes)"]
VMM["VMM: kvm, lxc"]
TM["TM / DS: ssh, shared, local, ceph, lvm, netapp, purefa"]
VNM["VNM: bridge, 802.1Q, VXLAN, Open vSwitch"]
AUTH["AUTH: core, ssh, x509, ldap, saml"]
IM["IM: monitoring probes"]
end
subgraph Hosts["Hypervisor hosts"]
LIBVIRT["libvirtd + QEMU/KVM"]
LXCH["LXC"]
end
FIREEDGE --> ONED
ONEFLOW --> ONED
ONEGATE --> ONED
ONEFORM --> ONED
ONEKS --> ONED
ONED --> DB
ONED --> SCHED
ONED --> HEM
MON --> ONED
ONED --> Drivers
VMM -->|"SSH"| LIBVIRT
VMM -->|"SSH"| LXCH
IM -->|"probes push to :4124"| MON
Core Daemon: oned¶
oned is installed by the opennebula package and runs as the opennebula systemd service. It:
- Manages hosts, clusters, virtual networks, datastores, images, templates, users, groups, VDCs and zones.
- Owns the VM lifecycle state machine (see VM Lifecycle).
- Exposes the XML-RPC API on port 2633 and, since 7.2, a gRPC API on port 2634 (enabled by default).
- Publishes events on a ZeroMQ socket (default
tcp://localhost:2101) consumed by hooks, OneKS and other services. - Calls drivers for every infrastructure action and stores results in the database.
Configuration lives in /etc/one/oned.conf. Because oned is the single source of truth, troubleshooting usually starts and ends with /var/log/one/oned.log.
Front-end High Availability and Federation¶
- HA: three or more front-ends run
onedreplicated with a built-in Raft consensus implementation; one leader serves the API behind a floating IP, followers replicate the database log. gRPC clients needENDPOINT_GRPCset on every HA server. - Federation: several independent OpenNebula instances (zones) share users, groups and ACLs; each zone keeps its own hosts and VMs. Federation is the recommended way past the certified scale of a single instance (500 hypervisors).
Driver Framework¶
Every infrastructure action is a script under /var/lib/one/remotes/ (synced to hosts with onehost sync). The driver types are:
| Driver type | Responsibility |
|---|---|
| VMM (virtual machine manager) | Deploy, shut down, migrate, snapshot, attach and detach on the hypervisor (kvm, lxc) |
| TM (transfer manager) | Move disk images between image and system datastores (copy over SSH, clone Ceph RBD, create LVs) |
| DS (datastore) | Register, clone, and delete images inside a datastore backend |
| VNM (virtual network manager) | Create bridges, VLANs, VXLAN tunnels, OVS flows, security-group rules on the host |
| IM (information manager) | Monitoring probes that push host and VM metrics to the monitor daemon |
| AUTH | Validate credentials (core, ssh, x509, ldap, saml) |
| MARKET, IPAM, backup | Marketplaces, external IP management, Restic/rsync/Veeam backups |
This design is why OpenNebula integrates a new storage array or network fabric by adding a driver rather than a new service. It also explains the security model: the front-end's oneadmin user must reach every host over SSH.
The full driver lists are in Reference: Drivers.
Hypervisors: KVM and LXC¶
OpenNebula 7.x supports exactly two virtualization drivers:
- KVM (QEMU via libvirt) is the primary hypervisor and carries almost all advanced features: live migration, live storage migration (7.2), memory encryption and vTPM (7.2), GPU and PCI passthrough with NUMA-aware topology, NVIDIA vGPU and NVLink/NVSwitch.
- LXC runs system containers with VM-like management. Containers are unprivileged by default. Some VM actions are not implemented for LXC: live migration, live disk resize, state save/restore, system snapshots, live qcow2 disk snapshots, live disk save and live capacity resize. 7.2 added NIC hot-plug, recontextualization, NIC PCI passthrough and disk snapshots on LVM, LVM thin, Ceph RBD and raw images.
Firecracker and vCenter are gone
The Firecracker microVM driver and the LXD driver were removed in 6.10 (alternatives: KVM and LXC). The vCenter driver was a legacy component in 6.10 and was removed in 7.0; VMware estates are now migrated with OneSwap. Older material that lists Firecracker or vCenter as supported refers to 6.8 or earlier.
Scheduler and OneDRS¶
The 7.x Scheduler Framework¶
OpenNebula 7.0 rewrote scheduling. The old persistent mm_sched daemon was replaced by a driver framework:
- Scheduler Manager inside
onedcollects placement requests during a scheduling window and asks for a plan. - Schedulers are external drivers (
one_sched) that compute a plan:rank(the classic matchmaking algorithm) orone_drs. - Scheduler Plan Manager executes the plan by deploying or migrating VMs, throttled by
MAX_ACTIONS_PER_HOSTandMAX_ACTIONS_PER_CLUSTER.
The scheduler is no longer a running service; it is executed on demand when a plan is needed. By default rank handles initial placement and OneDRS handles optimization (SCHED_MAD arguments -p rank -o one_drs).
It allocates three resource types: hosts, system datastores, and virtual networks (for NICs with NETWORK_MODE = "auto"). Placement combines automatic requirements (for example, a disk from a datastore in cluster 100 adds CLUSTER_ID = 100) with user expressions such as SCHED_REQUIREMENTS = "QOS != GOLD & HYPERVISOR = kvm".
OneDRS¶
OneDRS (OpenNebula Distributed Resource Scheduler, introduced in 7.0) balances workload inside a single Cluster:
- It uses an integer linear programming (ILP) solver, not a machine-learning model. The "predictive" part comes from the 7.0 resource-usage forecasting module, which projects CPU, memory, disk and network usage from monitoring history.
- Automation levels:
manual(recommendations only),partial(periodic plans that need approval),full(applied automatically). - Policies: packing (fewer active hosts, for energy or maintenance) or load balancing with weights for CPU usage, CPU allocation, memory, disk I/O and network.
- A VM is moved either to another host or to another datastore in a plan, never both at once. Since 7.4 OneDRS can also plan storage migrations, prioritize them, and skip VMs marked
ONEDRS_BLOCKED = "YES".
The flow below shows how a placement request and a OneDRS optimization cycle move through the framework.
sequenceDiagram
participant U as Sunstone or CLI
participant O as oned
participant SM as Scheduler Manager
participant S as one_sched (rank or one_drs)
participant PM as Plan Manager
participant H as KVM host (libvirt)
U->>O: one.template.instantiate
O->>O: VM PENDING, stored in DB
O->>SM: add VM to scheduling window
SM->>S: request placement plan
S-->>SM: plan: VM to host and system DS
SM->>PM: execute plan
PM->>O: deploy action
O->>H: TM prolog (copy or clone disks) over SSH
O->>H: VMM deploy (virsh create)
H-->>O: monitor reports RUNNING
Note over S,PM: OneDRS optimization cycle uses forecasts and emits migrate actions under the same throttles
Service Layer¶
Sunstone (FireEdge)¶
Sunstone is the web UI, served by the Node.js FireEdge server on port 2616 (/fireedge/sunstone). The Ruby Sunstone was deprecated in 6.10 and removed in 7.0. Since 7.2 the Node.js 20 runtime ships in the OpenNebula repositories. 7.2 added integrated VM logs and global two-factor authentication enforcement. 7.4 is a full visual redesign (new design system, side-drawer resource exploration, VirtioFS file-system management) with the same underlying concepts and workflows. FireEdge also brokers VNC, RDP and SSH consoles through Guacamole (opennebula-guacd).
OneFlow and OneGate¶
- OneFlow (port 2474) defines multi-VM services as roles with cardinality, deployment order and elasticity rules. Since 7.0 roles can also be Virtual Routers (
type: vr), and the service data model renamedvm_templatetotemplate_idandcustom_attrstouser_inputs. OneFlow can talk toonedover gRPC. - OneGate (port 5030) lets processes inside a VM read their own context and push metrics or attributes back, for example to drive OneFlow elasticity. VMs reach it directly or through the transparent proxy (TProxy) on the hypervisor.
OneForm (Cluster Provisioning)¶
OneForm (7.2) replaces the old oneprovision hybrid drivers. A Provider stores credentials and endpoints for a cloud or bare-metal source (for example AWS, i3D.net, Scaleway, or on-premises hardware). A Provision instantiates a cluster from a template: OneForm runs Terraform to create the machines and Ansible (OneDeploy roles) to install hypervisor software and register them with the front-end. CLI tools are oneform, oneprovider and oneprovision, also usable from Sunstone and a REST API.
Kubernetes: OneKE and OneKS¶
OpenNebula offers two Kubernetes paths:
- OneKE (OpenNebula Kubernetes Engine) is a marketplace service appliance, based on RKE2, deployed as a OneFlow service with control-plane, worker, storage (Longhorn) and virtual-router roles. The one-apps project still ships it, along with Cluster API work (CAPONE provider and a Rancher CAPI appliance).
- OneKS (Elastic Kubernetes as a Service) is a front-end service (
opennebula-ks, API on port 10780) that creates, scales, upgrades, recovers and deprovisions Kubernetes clusters as a managed service. It keeps its own RKE2-based management context on the front-end (kubectlfrom/var/lib/rancher/rke2) and deploys clusters from a marketplace appliance. It shipped to EE customers in 7.2.1 and became a Community Edition feature in 7.4, which also added multi-cluster placement and pre-deployment readiness checks. Certified Kubernetes versions in 7.4 are 1.33.7 and 1.34.2.
Storage Model¶
OpenNebula separates storage into three datastore roles:
| Role | Holds | Example path |
|---|---|---|
| Image datastore | Registered images (OS disks, ISOs, persistent data disks) | /var/lib/one/datastores/1/ |
| System datastore | Running VMs' disks, checkpoints, context ISO | /var/lib/one/datastores/0/ |
| File datastore | Kernels, initrds, context files | /var/lib/one/datastores/2/ |
| Backup datastore | Restic or rsync backup repositories | Backup server or S3 bucket (Restic S3 in 7.4) |
The transfer driver decides how an image becomes a running disk:
| Backend | Access pattern | Trade-off |
|---|---|---|
| Shared FS (NFS) | Front-end and hosts mount the same export | Simple; live migration without copying disks |
| SSH / local | Images copied to host-local disks, with optional multi-tier caching | No shared storage needed; migration copies disks |
| Ceph RBD | Hosts access RBD images directly; clones are copy-on-write | Scales well; live migration native |
| LVM / LVM thin on SAN (iSCSI, FC) | Logical volumes on shared LUNs | Block performance; redesigned in 7.2 (EE), CE since 7.4 |
| NetApp ONTAP, Pure Storage FlashArray | Array-native drivers over the vendor REST API | Offload snapshots and clones to the array |
| VirtioFS | Host file systems shared into VMs | Low-latency shared file access (7.2) |
7.2 added live storage migration between LVM and file-based datastores. 7.4 added OneBEX, an exporter that lets third-party backup tools pull full and incremental qcow2 and LVM disk data straight from hypervisors. Veeam now uses it instead of a separate backup-server VM.
Networking Model¶
A Virtual Network (VNet) couples a network driver with one or more Address Ranges (IPv4, IPv6, dual-stack or Ethernet-only). The driver builds the host-side plumbing when a NIC is attached.
This diagram shows the host data path for a VM NIC with security groups.
flowchart LR
subgraph VM["Virtual Machine"]
VNIC["virtio NIC"]
end
subgraph Host["KVM host"]
TAP["tap device"]
SG["Security group rules<br/>(iptables/ipset on the host)"]
BR["Linux bridge or OVS bridge"]
ENC["802.1Q VLAN or VXLAN VNI"]
PHY["Physical NIC or SR-IOV VF"]
end
VNIC --> TAP --> SG --> BR --> ENC --> PHY
Other networking features: Virtual Routers (appliance VMs for routing, NAT and DHCP), shared Address Ranges with NIC aliases for virtual IPs (7.2), DPDK vhost-user NICs, SR-IOV switchdev with Open vSwitch (7.4), VLAN delegation rules that let tenants self-provision networks (7.4), and a round-robin lease policy that avoids immediately reusing freed addresses (7.4).
VM Lifecycle¶
The state diagram shows the main OpenNebula VM states and the CLI actions that move between them.
stateDiagram-v2
[*] --> PENDING : onetemplate instantiate
PENDING --> HOLD : onevm hold
HOLD --> PENDING : onevm release
PENDING --> PROLOG : scheduler deploys
PROLOG --> BOOT : disks transferred
BOOT --> RUNNING : hypervisor started guest
RUNNING --> MIGRATE : onevm migrate
MIGRATE --> RUNNING : migration done
RUNNING --> SUSPENDED : onevm suspend
SUSPENDED --> RUNNING : onevm resume
RUNNING --> POWEROFF : onevm poweroff
POWEROFF --> RUNNING : onevm resume
RUNNING --> UNDEPLOYED : onevm undeploy
UNDEPLOYED --> PENDING : onevm resume
RUNNING --> EPILOG : onevm terminate
POWEROFF --> EPILOG : onevm terminate
EPILOG --> DONE : cleanup done
PROLOG --> FAILURE : transfer error
BOOT --> FAILURE : boot error
RUNNING --> UNKNOWN : host unreachable
DONE --> [*]
RUNNING, PROLOG, BOOT, MIGRATE and EPILOG are sub-states (LCM states) of the ACTIVE state; the CLI shows the short forms runn, prol, boot, migr, epil. The full state table is in Reference: VM States.
Security Model¶
Identity and Authentication¶
Each user has one authentication driver. Different drivers can be enabled at the same time and chosen per user:
| Driver | Method | Usable from |
|---|---|---|
core |
Username/password (hashed in DB) plus login tokens | API, CLI, Sunstone |
ssh |
SSH key-pair challenge, generating a login token | API, CLI |
x509 |
Client certificate | API, CLI |
ldap |
Bind against LDAP / Active Directory | API, CLI, Sunstone |
saml |
SAML 2.0 federation (added in 7.0.1, off by default) | Sunstone |
server_cipher, server_x509 |
Service accounts for servers such as FireEdge and OneFlow that act on behalf of users | Internal |
OpenID Connect
The 7.4 authentication overview lists no native OIDC driver, and the source tree agrees: the default AUTH_MAD list in oned.conf is ssh,x509,ldap,server_cipher,server_x509,saml, and install.sh ships no OIDC driver (oned.conf on master, checked 2026-09-27). OIDC single sign-on has to go through Sunstone's remote/external authentication behind a reverse proxy.
The diagram shows how requests are authenticated and then authorized.
graph LR
subgraph Clients
SUN["Sunstone user"]
CLI["CLI / pyone / Go client"]
SVC["OneFlow, FireEdge (server_cipher)"]
end
subgraph oned["oned"]
AM["Auth Manager<br/>(core, ssh, x509, ldap, saml drivers)"]
ACL["ACL engine + permission bits"]
QUOTA["Quotas<br/>(user, group, cluster-level)"]
end
SUN --> AM
CLI --> AM
SVC -->|"acts on behalf of user"| AM
AM --> ACL --> QUOTA
Authorization, Tenancy and Quotas¶
Authorization combines Unix-like permission bits on each object (owner, group, other x use, manage, admin) with ACL rules (oneacl) that grant rights on resource types to users or groups. Virtual Data Centers (VDCs) group one or more user groups with a set of clusters, hosts, datastores and virtual networks from one or more zones. The same physical resource can be assigned to more than one VDC, so a VDC is an administrative boundary for delegation and self-service, not a hard isolation boundary. Hard isolation comes from separate clusters, VLAN/VXLAN segmentation, security groups and host affinity.
7.0 added cluster-level quotas (per user or group per cluster, useful at the edge) and generic quotas for custom resources such as vGPUs or licenses.
Workload Isolation¶
| Mechanism | What it isolates |
|---|---|
| KVM hardware virtualization + libvirt sVirt (SELinux/AppArmor on the host) | Guest from host and other guests |
| VM memory encryption (7.2; needs CPU memory-encryption support on the host) | Guest memory from the hypervisor |
| vTPM (7.0.x/7.2) | Measured boot and guest secrets |
| LXC unprivileged containers (UID/GID mapping to 600100001+) | Container root from host root |
VM Groups affinity / anti-affinity, host pinning via SCHED_REQUIREMENTS |
Placement for fault tolerance and dedicated tenancy |
| Security groups | L3/L4 traffic per NIC |
Threat Model Notes¶
- The front-end is the crown jewel:
oneadminthere can SSH to every host and read every credential in the DB. Protect it like a hypervisor manager. - The API ports (2633, 2634) and OneGate should not be exposed to untrusted networks. Sunstone should sit behind TLS.
- Drivers run as
oneadminon hosts over passwordless SSH (keys held byopennebula-ssh-agenton the front-end), so a compromised host is a foothold into cloud operations. Keep host access restricted and patched.
The concrete checklist is in Reference: Hardening Checklist.
Performance and Scale¶
Official Numbers¶
- A single
onedinstance is certified for 500 hypervisors without degradation; users run up to about 2,000 per instance, and the key-features page cites over 2,500 nodes in production. Past 500, the recommendation is federation. - The 7.2 gRPC API cut average
vmpool.infolatency from 0.97 s to 0.43 s at 10 req/s and from 2.15 s to 0.94 s at 30 req/s in a synthetic 1,250-host / 20,000-VM test (see Reference: Published Benchmarks). The saving comes from binary Protocol Buffers instead of XML serialization. The XML-RPC API remains available and is still used by several services by default.
Hypervisor Overhead and Scheduler Timing¶
OpenNebula publishes no hypervisor-overhead or scheduler-latency benchmarks; the only official scale figures are the ones above (checked 2026-09-27). Earlier unsourced estimates in this note (KVM/LXC overhead percentages, scheduler timings per 10/100/1,000 VMs) were removed. Measure on your own hardware before capacity planning.
Why OpenNebula Is Shaped This Way¶
Centralized by Design¶
The single-daemon design trades horizontal control-plane scale-out for operational simplicity: one process, one database, one log, and upgrades that are mostly "stop, upgrade packages, onedb upgrade, start". Scale beyond one instance is handled by federation and many small edge clusters rather than by splitting the control plane into services. This fits OpenNebula's target of many distributed clusters (edge sites, provider regions) managed from one place.
From VMware Alternative to AI Factory Platform¶
The 6.10 and 7.0 releases removed the vCenter driver, Ruby Sunstone, Firecracker, LXD and old hybrid drivers to concentrate on KVM and LXC. That clean-up lined up with Broadcom's VMware licensing changes: OneSwap (6.10, batch mode in 7.4) converts vCenter VMs and OVAs to KVM, and OneDRS (7.0) fills the gap left by VMware DRS. 7.2 and 7.4 then pushed into "AI factories": NVIDIA Fabric Manager for NVLink/NVSwitch, Grace Blackwell GB200/GB300 validation, Spectrum-X and BlueField DPU support, the Slurm appliance, and in 7.4 NVIDIA Infra Controller (NICo) integration for bare-metal as a service (EE).
European Sovereign-Cloud Context (IPCEI-CIS)¶
OpenNebula Systems (Madrid, Spain) is a core participant in IPCEI-CIS (Important Project of Common European Interest on Next Generation Cloud Infrastructure and Services), approved by the European Commission in December 2023 with about EUR 1.2 billion in state aid and EUR 1.4 billion in private investment. Its ONEnextgen project (UNICO IPCEI-2023-003, 2024-2028) is funded by the Spanish Ministry for Digital Transformation and Civil Service and co-funded by the EU NextGenerationEU Recovery and Resilience Facility. The 7.4 release notes credit ONEnextgen for part of the new functionality. OpenNebula Systems also led the IPCEI-CIS reference architecture work, and OpenNebula is the virtualization layer of Virt8ra, a multi-provider European sovereign cloud initiative. A separate project, ONEedge5G (TSI-064200-2023-1), funded edge and 5G work. This funding explains the roadmap's emphasis on cloud-edge continuum, multi-provider federation, and sovereignty.
Sources¶
- OpenNebula 7.4 docs (scheduler overview, OneDRS, gRPC integration, OneKS configuration, OneForm overview, LXC driver, authentication overview, front-end installation), read from the
OpenNebula/websitedocs source repository - Release Notes 7.4 - What's New
- Release Notes 7.2
- 7.0 Compatibility Guide
- OneKE Service docs (7.0)
- Introducing OneKS
- IPCEI-CIS initiative page and Virt8ra providers announcement