Skip to content

Explanation

Scope

How OpenNebula works and why it is built that way: the centralized oned control plane, the driver framework, the 7.x scheduler and OneDRS, the service layer (Sunstone, OneFlow, OneGate, OneForm, OneKS), storage and networking models, the security model, performance characteristics, and the European sovereign-cloud context behind the 7.x roadmap. Look-up tables (ports, drivers, config keys, VM states) live in Reference; commands live in How-to Guides.

OpenNebula is a centralized cloud management platform. Its core runs as one C++ daemon (oned) on a front-end node, which keeps all state in a SQL database and drives hypervisor hosts, storage and networks through pluggable driver scripts executed over SSH. This is the opposite of OpenStack's design of many cooperating services, and it is the main reason OpenNebula is considered simpler to install and operate.

See also: OpenNebula hub, Reference, How-to Guides.

Architecture

Component Overview

The diagram shows the 7.4 front-end services, the driver framework, and what each driver talks to on the hosts.

graph TD
    subgraph Frontend["Front-end node"]
        ONED["oned<br/>(core daemon, XML-RPC :2633, gRPC :2634)"]
        DB[("SQLite or MariaDB/MySQL")]
        SCHED["one_sched driver<br/>(rank + OneDRS)"]
        MON["Monitor daemon<br/>(:4124)"]
        HEM["Hook Execution Manager"]
        FIREEDGE["FireEdge / Sunstone<br/>(:2616)"]
        ONEFLOW["OneFlow (:2474)"]
        ONEGATE["OneGate (:5030)"]
        ONEFORM["OneForm<br/>(cluster provisioning)"]
        ONEKS["OneKS<br/>(Kubernetes as a Service)"]
    end

    subgraph Drivers["Driver framework (/var/lib/one/remotes)"]
        VMM["VMM: kvm, lxc"]
        TM["TM / DS: ssh, shared, local, ceph, lvm, netapp, purefa"]
        VNM["VNM: bridge, 802.1Q, VXLAN, Open vSwitch"]
        AUTH["AUTH: core, ssh, x509, ldap, saml"]
        IM["IM: monitoring probes"]
    end

    subgraph Hosts["Hypervisor hosts"]
        LIBVIRT["libvirtd + QEMU/KVM"]
        LXCH["LXC"]
    end

    FIREEDGE --> ONED
    ONEFLOW --> ONED
    ONEGATE --> ONED
    ONEFORM --> ONED
    ONEKS --> ONED
    ONED --> DB
    ONED --> SCHED
    ONED --> HEM
    MON --> ONED
    ONED --> Drivers
    VMM -->|"SSH"| LIBVIRT
    VMM -->|"SSH"| LXCH
    IM -->|"probes push to :4124"| MON

Core Daemon: oned

oned is installed by the opennebula package and runs as the opennebula systemd service. It:

  • Manages hosts, clusters, virtual networks, datastores, images, templates, users, groups, VDCs and zones.
  • Owns the VM lifecycle state machine (see VM Lifecycle).
  • Exposes the XML-RPC API on port 2633 and, since 7.2, a gRPC API on port 2634 (enabled by default).
  • Publishes events on a ZeroMQ socket (default tcp://localhost:2101) consumed by hooks, OneKS and other services.
  • Calls drivers for every infrastructure action and stores results in the database.

Configuration lives in /etc/one/oned.conf. Because oned is the single source of truth, troubleshooting usually starts and ends with /var/log/one/oned.log.

Front-end High Availability and Federation

  • HA: three or more front-ends run oned replicated with a built-in Raft consensus implementation; one leader serves the API behind a floating IP, followers replicate the database log. gRPC clients need ENDPOINT_GRPC set on every HA server.
  • Federation: several independent OpenNebula instances (zones) share users, groups and ACLs; each zone keeps its own hosts and VMs. Federation is the recommended way past the certified scale of a single instance (500 hypervisors).

Driver Framework

Every infrastructure action is a script under /var/lib/one/remotes/ (synced to hosts with onehost sync). The driver types are:

Driver type Responsibility
VMM (virtual machine manager) Deploy, shut down, migrate, snapshot, attach and detach on the hypervisor (kvm, lxc)
TM (transfer manager) Move disk images between image and system datastores (copy over SSH, clone Ceph RBD, create LVs)
DS (datastore) Register, clone, and delete images inside a datastore backend
VNM (virtual network manager) Create bridges, VLANs, VXLAN tunnels, OVS flows, security-group rules on the host
IM (information manager) Monitoring probes that push host and VM metrics to the monitor daemon
AUTH Validate credentials (core, ssh, x509, ldap, saml)
MARKET, IPAM, backup Marketplaces, external IP management, Restic/rsync/Veeam backups

This design is why OpenNebula integrates a new storage array or network fabric by adding a driver rather than a new service. It also explains the security model: the front-end's oneadmin user must reach every host over SSH.

The full driver lists are in Reference: Drivers.

Hypervisors: KVM and LXC

OpenNebula 7.x supports exactly two virtualization drivers:

  • KVM (QEMU via libvirt) is the primary hypervisor and carries almost all advanced features: live migration, live storage migration (7.2), memory encryption and vTPM (7.2), GPU and PCI passthrough with NUMA-aware topology, NVIDIA vGPU and NVLink/NVSwitch.
  • LXC runs system containers with VM-like management. Containers are unprivileged by default. Some VM actions are not implemented for LXC: live migration, live disk resize, state save/restore, system snapshots, live qcow2 disk snapshots, live disk save and live capacity resize. 7.2 added NIC hot-plug, recontextualization, NIC PCI passthrough and disk snapshots on LVM, LVM thin, Ceph RBD and raw images.

Firecracker and vCenter are gone

The Firecracker microVM driver and the LXD driver were removed in 6.10 (alternatives: KVM and LXC). The vCenter driver was a legacy component in 6.10 and was removed in 7.0; VMware estates are now migrated with OneSwap. Older material that lists Firecracker or vCenter as supported refers to 6.8 or earlier.

Scheduler and OneDRS

The 7.x Scheduler Framework

OpenNebula 7.0 rewrote scheduling. The old persistent mm_sched daemon was replaced by a driver framework:

  1. Scheduler Manager inside oned collects placement requests during a scheduling window and asks for a plan.
  2. Schedulers are external drivers (one_sched) that compute a plan: rank (the classic matchmaking algorithm) or one_drs.
  3. Scheduler Plan Manager executes the plan by deploying or migrating VMs, throttled by MAX_ACTIONS_PER_HOST and MAX_ACTIONS_PER_CLUSTER.

The scheduler is no longer a running service; it is executed on demand when a plan is needed. By default rank handles initial placement and OneDRS handles optimization (SCHED_MAD arguments -p rank -o one_drs).

It allocates three resource types: hosts, system datastores, and virtual networks (for NICs with NETWORK_MODE = "auto"). Placement combines automatic requirements (for example, a disk from a datastore in cluster 100 adds CLUSTER_ID = 100) with user expressions such as SCHED_REQUIREMENTS = "QOS != GOLD & HYPERVISOR = kvm".

OneDRS

OneDRS (OpenNebula Distributed Resource Scheduler, introduced in 7.0) balances workload inside a single Cluster:

  • It uses an integer linear programming (ILP) solver, not a machine-learning model. The "predictive" part comes from the 7.0 resource-usage forecasting module, which projects CPU, memory, disk and network usage from monitoring history.
  • Automation levels: manual (recommendations only), partial (periodic plans that need approval), full (applied automatically).
  • Policies: packing (fewer active hosts, for energy or maintenance) or load balancing with weights for CPU usage, CPU allocation, memory, disk I/O and network.
  • A VM is moved either to another host or to another datastore in a plan, never both at once. Since 7.4 OneDRS can also plan storage migrations, prioritize them, and skip VMs marked ONEDRS_BLOCKED = "YES".

The flow below shows how a placement request and a OneDRS optimization cycle move through the framework.

sequenceDiagram
    participant U as Sunstone or CLI
    participant O as oned
    participant SM as Scheduler Manager
    participant S as one_sched (rank or one_drs)
    participant PM as Plan Manager
    participant H as KVM host (libvirt)

    U->>O: one.template.instantiate
    O->>O: VM PENDING, stored in DB
    O->>SM: add VM to scheduling window
    SM->>S: request placement plan
    S-->>SM: plan: VM to host and system DS
    SM->>PM: execute plan
    PM->>O: deploy action
    O->>H: TM prolog (copy or clone disks) over SSH
    O->>H: VMM deploy (virsh create)
    H-->>O: monitor reports RUNNING
    Note over S,PM: OneDRS optimization cycle uses forecasts and emits migrate actions under the same throttles

Service Layer

Sunstone (FireEdge)

Sunstone is the web UI, served by the Node.js FireEdge server on port 2616 (/fireedge/sunstone). The Ruby Sunstone was deprecated in 6.10 and removed in 7.0. Since 7.2 the Node.js 20 runtime ships in the OpenNebula repositories. 7.2 added integrated VM logs and global two-factor authentication enforcement. 7.4 is a full visual redesign (new design system, side-drawer resource exploration, VirtioFS file-system management) with the same underlying concepts and workflows. FireEdge also brokers VNC, RDP and SSH consoles through Guacamole (opennebula-guacd).

OneFlow and OneGate

  • OneFlow (port 2474) defines multi-VM services as roles with cardinality, deployment order and elasticity rules. Since 7.0 roles can also be Virtual Routers (type: vr), and the service data model renamed vm_template to template_id and custom_attrs to user_inputs. OneFlow can talk to oned over gRPC.
  • OneGate (port 5030) lets processes inside a VM read their own context and push metrics or attributes back, for example to drive OneFlow elasticity. VMs reach it directly or through the transparent proxy (TProxy) on the hypervisor.

OneForm (Cluster Provisioning)

OneForm (7.2) replaces the old oneprovision hybrid drivers. A Provider stores credentials and endpoints for a cloud or bare-metal source (for example AWS, i3D.net, Scaleway, or on-premises hardware). A Provision instantiates a cluster from a template: OneForm runs Terraform to create the machines and Ansible (OneDeploy roles) to install hypervisor software and register them with the front-end. CLI tools are oneform, oneprovider and oneprovision, also usable from Sunstone and a REST API.

Kubernetes: OneKE and OneKS

OpenNebula offers two Kubernetes paths:

  • OneKE (OpenNebula Kubernetes Engine) is a marketplace service appliance, based on RKE2, deployed as a OneFlow service with control-plane, worker, storage (Longhorn) and virtual-router roles. The one-apps project still ships it, along with Cluster API work (CAPONE provider and a Rancher CAPI appliance).
  • OneKS (Elastic Kubernetes as a Service) is a front-end service (opennebula-ks, API on port 10780) that creates, scales, upgrades, recovers and deprovisions Kubernetes clusters as a managed service. It keeps its own RKE2-based management context on the front-end (kubectl from /var/lib/rancher/rke2) and deploys clusters from a marketplace appliance. It shipped to EE customers in 7.2.1 and became a Community Edition feature in 7.4, which also added multi-cluster placement and pre-deployment readiness checks. Certified Kubernetes versions in 7.4 are 1.33.7 and 1.34.2.

Storage Model

OpenNebula separates storage into three datastore roles:

Role Holds Example path
Image datastore Registered images (OS disks, ISOs, persistent data disks) /var/lib/one/datastores/1/
System datastore Running VMs' disks, checkpoints, context ISO /var/lib/one/datastores/0/
File datastore Kernels, initrds, context files /var/lib/one/datastores/2/
Backup datastore Restic or rsync backup repositories Backup server or S3 bucket (Restic S3 in 7.4)

The transfer driver decides how an image becomes a running disk:

Backend Access pattern Trade-off
Shared FS (NFS) Front-end and hosts mount the same export Simple; live migration without copying disks
SSH / local Images copied to host-local disks, with optional multi-tier caching No shared storage needed; migration copies disks
Ceph RBD Hosts access RBD images directly; clones are copy-on-write Scales well; live migration native
LVM / LVM thin on SAN (iSCSI, FC) Logical volumes on shared LUNs Block performance; redesigned in 7.2 (EE), CE since 7.4
NetApp ONTAP, Pure Storage FlashArray Array-native drivers over the vendor REST API Offload snapshots and clones to the array
VirtioFS Host file systems shared into VMs Low-latency shared file access (7.2)

7.2 added live storage migration between LVM and file-based datastores. 7.4 added OneBEX, an exporter that lets third-party backup tools pull full and incremental qcow2 and LVM disk data straight from hypervisors. Veeam now uses it instead of a separate backup-server VM.

Networking Model

A Virtual Network (VNet) couples a network driver with one or more Address Ranges (IPv4, IPv6, dual-stack or Ethernet-only). The driver builds the host-side plumbing when a NIC is attached.

This diagram shows the host data path for a VM NIC with security groups.

flowchart LR
    subgraph VM["Virtual Machine"]
        VNIC["virtio NIC"]
    end
    subgraph Host["KVM host"]
        TAP["tap device"]
        SG["Security group rules<br/>(iptables/ipset on the host)"]
        BR["Linux bridge or OVS bridge"]
        ENC["802.1Q VLAN or VXLAN VNI"]
        PHY["Physical NIC or SR-IOV VF"]
    end
    VNIC --> TAP --> SG --> BR --> ENC --> PHY

Other networking features: Virtual Routers (appliance VMs for routing, NAT and DHCP), shared Address Ranges with NIC aliases for virtual IPs (7.2), DPDK vhost-user NICs, SR-IOV switchdev with Open vSwitch (7.4), VLAN delegation rules that let tenants self-provision networks (7.4), and a round-robin lease policy that avoids immediately reusing freed addresses (7.4).

VM Lifecycle

The state diagram shows the main OpenNebula VM states and the CLI actions that move between them.

stateDiagram-v2
    [*] --> PENDING : onetemplate instantiate
    PENDING --> HOLD : onevm hold
    HOLD --> PENDING : onevm release
    PENDING --> PROLOG : scheduler deploys
    PROLOG --> BOOT : disks transferred
    BOOT --> RUNNING : hypervisor started guest
    RUNNING --> MIGRATE : onevm migrate
    MIGRATE --> RUNNING : migration done
    RUNNING --> SUSPENDED : onevm suspend
    SUSPENDED --> RUNNING : onevm resume
    RUNNING --> POWEROFF : onevm poweroff
    POWEROFF --> RUNNING : onevm resume
    RUNNING --> UNDEPLOYED : onevm undeploy
    UNDEPLOYED --> PENDING : onevm resume
    RUNNING --> EPILOG : onevm terminate
    POWEROFF --> EPILOG : onevm terminate
    EPILOG --> DONE : cleanup done
    PROLOG --> FAILURE : transfer error
    BOOT --> FAILURE : boot error
    RUNNING --> UNKNOWN : host unreachable
    DONE --> [*]

RUNNING, PROLOG, BOOT, MIGRATE and EPILOG are sub-states (LCM states) of the ACTIVE state; the CLI shows the short forms runn, prol, boot, migr, epil. The full state table is in Reference: VM States.

Security Model

Identity and Authentication

Each user has one authentication driver. Different drivers can be enabled at the same time and chosen per user:

Driver Method Usable from
core Username/password (hashed in DB) plus login tokens API, CLI, Sunstone
ssh SSH key-pair challenge, generating a login token API, CLI
x509 Client certificate API, CLI
ldap Bind against LDAP / Active Directory API, CLI, Sunstone
saml SAML 2.0 federation (added in 7.0.1, off by default) Sunstone
server_cipher, server_x509 Service accounts for servers such as FireEdge and OneFlow that act on behalf of users Internal

OpenID Connect

The 7.4 authentication overview lists no native OIDC driver, and the source tree agrees: the default AUTH_MAD list in oned.conf is ssh,x509,ldap,server_cipher,server_x509,saml, and install.sh ships no OIDC driver (oned.conf on master, checked 2026-09-27). OIDC single sign-on has to go through Sunstone's remote/external authentication behind a reverse proxy.

The diagram shows how requests are authenticated and then authorized.

graph LR
    subgraph Clients
        SUN["Sunstone user"]
        CLI["CLI / pyone / Go client"]
        SVC["OneFlow, FireEdge (server_cipher)"]
    end
    subgraph oned["oned"]
        AM["Auth Manager<br/>(core, ssh, x509, ldap, saml drivers)"]
        ACL["ACL engine + permission bits"]
        QUOTA["Quotas<br/>(user, group, cluster-level)"]
    end
    SUN --> AM
    CLI --> AM
    SVC -->|"acts on behalf of user"| AM
    AM --> ACL --> QUOTA

Authorization, Tenancy and Quotas

Authorization combines Unix-like permission bits on each object (owner, group, other x use, manage, admin) with ACL rules (oneacl) that grant rights on resource types to users or groups. Virtual Data Centers (VDCs) group one or more user groups with a set of clusters, hosts, datastores and virtual networks from one or more zones. The same physical resource can be assigned to more than one VDC, so a VDC is an administrative boundary for delegation and self-service, not a hard isolation boundary. Hard isolation comes from separate clusters, VLAN/VXLAN segmentation, security groups and host affinity.

7.0 added cluster-level quotas (per user or group per cluster, useful at the edge) and generic quotas for custom resources such as vGPUs or licenses.

Workload Isolation

Mechanism What it isolates
KVM hardware virtualization + libvirt sVirt (SELinux/AppArmor on the host) Guest from host and other guests
VM memory encryption (7.2; needs CPU memory-encryption support on the host) Guest memory from the hypervisor
vTPM (7.0.x/7.2) Measured boot and guest secrets
LXC unprivileged containers (UID/GID mapping to 600100001+) Container root from host root
VM Groups affinity / anti-affinity, host pinning via SCHED_REQUIREMENTS Placement for fault tolerance and dedicated tenancy
Security groups L3/L4 traffic per NIC

Threat Model Notes

  • The front-end is the crown jewel: oneadmin there can SSH to every host and read every credential in the DB. Protect it like a hypervisor manager.
  • The API ports (2633, 2634) and OneGate should not be exposed to untrusted networks. Sunstone should sit behind TLS.
  • Drivers run as oneadmin on hosts over passwordless SSH (keys held by opennebula-ssh-agent on the front-end), so a compromised host is a foothold into cloud operations. Keep host access restricted and patched.

The concrete checklist is in Reference: Hardening Checklist.

Performance and Scale

Official Numbers

  • A single oned instance is certified for 500 hypervisors without degradation; users run up to about 2,000 per instance, and the key-features page cites over 2,500 nodes in production. Past 500, the recommendation is federation.
  • The 7.2 gRPC API cut average vmpool.info latency from 0.97 s to 0.43 s at 10 req/s and from 2.15 s to 0.94 s at 30 req/s in a synthetic 1,250-host / 20,000-VM test (see Reference: Published Benchmarks). The saving comes from binary Protocol Buffers instead of XML serialization. The XML-RPC API remains available and is still used by several services by default.

Hypervisor Overhead and Scheduler Timing

OpenNebula publishes no hypervisor-overhead or scheduler-latency benchmarks; the only official scale figures are the ones above (checked 2026-09-27). Earlier unsourced estimates in this note (KVM/LXC overhead percentages, scheduler timings per 10/100/1,000 VMs) were removed. Measure on your own hardware before capacity planning.

Why OpenNebula Is Shaped This Way

Centralized by Design

The single-daemon design trades horizontal control-plane scale-out for operational simplicity: one process, one database, one log, and upgrades that are mostly "stop, upgrade packages, onedb upgrade, start". Scale beyond one instance is handled by federation and many small edge clusters rather than by splitting the control plane into services. This fits OpenNebula's target of many distributed clusters (edge sites, provider regions) managed from one place.

From VMware Alternative to AI Factory Platform

The 6.10 and 7.0 releases removed the vCenter driver, Ruby Sunstone, Firecracker, LXD and old hybrid drivers to concentrate on KVM and LXC. That clean-up lined up with Broadcom's VMware licensing changes: OneSwap (6.10, batch mode in 7.4) converts vCenter VMs and OVAs to KVM, and OneDRS (7.0) fills the gap left by VMware DRS. 7.2 and 7.4 then pushed into "AI factories": NVIDIA Fabric Manager for NVLink/NVSwitch, Grace Blackwell GB200/GB300 validation, Spectrum-X and BlueField DPU support, the Slurm appliance, and in 7.4 NVIDIA Infra Controller (NICo) integration for bare-metal as a service (EE).

European Sovereign-Cloud Context (IPCEI-CIS)

OpenNebula Systems (Madrid, Spain) is a core participant in IPCEI-CIS (Important Project of Common European Interest on Next Generation Cloud Infrastructure and Services), approved by the European Commission in December 2023 with about EUR 1.2 billion in state aid and EUR 1.4 billion in private investment. Its ONEnextgen project (UNICO IPCEI-2023-003, 2024-2028) is funded by the Spanish Ministry for Digital Transformation and Civil Service and co-funded by the EU NextGenerationEU Recovery and Resilience Facility. The 7.4 release notes credit ONEnextgen for part of the new functionality. OpenNebula Systems also led the IPCEI-CIS reference architecture work, and OpenNebula is the virtualization layer of Virt8ra, a multi-provider European sovereign cloud initiative. A separate project, ONEedge5G (TSI-064200-2023-1), funded edge and 5G work. This funding explains the roadmap's emphasis on cloud-edge continuum, multi-provider federation, and sovereignty.

Sources