Longhorn¶
Summary
Longhorn is a lightweight, Kubernetes-native distributed block storage system. It is a CNCF Incubating project, originally built by Rancher Labs and now maintained mainly by SUSE. Every volume gets its own storage controller (engine) with synchronous replicas on other nodes. It is delivered through CSI and ships with incremental snapshots, backups to S3/NFS/SMB/Azure, cross-cluster DR volumes, system backup, and RWX through NFS share managers. v1.12.0 (June 2026) made the SPDK-based V2 Data Engine generally available beside the default iSCSI-based V1 engine. The latest release is v1.12.1 (2026-08-14).
Overview¶
Longhorn runs entirely inside Kubernetes. A longhorn-manager DaemonSet reconciles Longhorn CRs. Per-node instance-manager pods host the per-volume engines and replicas. The CSI driver driver.longhorn.io connects PVCs to volumes. Because it needs only local disks on worker nodes and a Helm chart, it is usually the simplest replicated block storage to run on small and medium clusters, edge sites (K3s, RKE2), and SUSE's Harvester HCI, which uses Longhorn as its storage layer.
Key Facts¶
| Attribute | Detail |
|---|---|
| Repository | github.com/longhorn/longhorn |
| Stars | ~6k+ (recorded by 2026-08) |
| Latest Version | v1.12.1 (2026-08-14). Also maintained: v1.11.3 (2026-07-02) |
| Release cadence | Minor release every ~4-5 months. Each branch reaches EOL one year after its first stable version. Upgrades must go one minor at a time |
| Language | Go (manager, V1 engine, share manager). V2 builds on SPDK (C) |
| License | Apache 2.0 |
| Governance | CNCF Incubating (Sandbox 2019, Incubating since 2021-11-04) |
| Company | SUSE (acquired Rancher Labs in 2020). Commercial support as SUSE Storage in Rancher Prime |
| Kubernetes | >= v1.25 (tested on 1.33 to 1.36 for v1.12.1) |
| Data engines | V1 (iSCSI, default, GA) and V2 (SPDK with NVMe-TCP frontend, GA since v1.12.0) |
| Access modes | RWO, RWOP (v1.11+), RWX (NFSv4.1 share manager), migratable RWX Block (KubeVirt) |
| Backup targets | S3-compatible (AWS, MinIO, GCS through its S3 API), NFS, SMB/CIFS, Azure Blob. Several targets since v1.8 |
Architecture at a Glance¶
A compact view of the data path: the CSI driver asks the manager to attach a volume, and the engine on the Pod's node replicates writes to replicas on other nodes.
flowchart LR
Pod["Pod + PVC"] --> CSI["longhorn-csi-plugin"]
CSI -->|"attach"| LHM["longhorn-manager<br/>(DaemonSet)"]
LHM --> IM1
subgraph N1["Node 1"]
IM1["instance-manager<br/>Engine (V1 iSCSI / V2 SPDK)"]
R1["Replica 1"]
end
subgraph N2["Node 2"]
R2["Replica 2"]
end
subgraph N3["Node 3"]
R3["Replica 3"]
end
IM1 -->|"sync write"| R1
IM1 -->|"sync write"| R2
IM1 -->|"sync write"| R3
IM1 -.->|"incremental 2 MiB blocks"| BS[("Backupstore<br/>S3 / NFS / SMB / Azure")]
Details are in Explanation: Architecture.
Evaluation¶
| Pros | Cons |
|---|---|
| Simplest replicated block storage for Kubernetes: Helm install, local disks only | Kubernetes-only. Needs privileged root pods and host packages (open-iscsi, NFS client) |
| Per-volume engine: small failure domains, per-volume replica count and locality | Write performance is bounded by synchronous replication over the network (see benchmarks) |
| Incremental, deduplicated backups to S3/NFS/SMB/Azure. DR standby volumes. System backup | Block storage only. RWX is NFS in front of a block volume (single share-manager pod per volume), not a distributed filesystem. No object storage |
| Live V1 engine upgrades. Guided upgrade checks | Strict one-minor-at-a-time upgrades. V2 volumes must be detached to upgrade |
| V2 Data Engine (SPDK) is GA for low-latency NVMe setups | V2 needs dedicated CPU cores, huge pages, raw block disks and kernel 6.7+. Some features are still V1-only |
Built-in UI, Prometheus metrics, longhornctl CLI |
The UI has no authentication of its own. Internal NetworkPolicies (v1.12.1) can surprise enforcing CNIs |
| CNCF Incubating, Apache 2.0, active SUSE-backed community | Control-plane throughput limits attach/detach at high volume counts. Official scale tests stop at 9 workers and about 1,500 volumes |
When Longhorn fits¶
- Small to medium Kubernetes clusters that need replicated RWO volumes without a storage team.
- Edge and on-prem clusters (K3s/RKE2), and Harvester/KubeVirt VM disks (migratable volumes).
- Teams that want backup/DR to object storage built into the storage layer.
When to look elsewhere¶
- You need unified block + object + file, or very large clusters: consider Ceph (through Rook).
- Your apps replicate their own data and need raw local-disk performance: consider local PVs, or Longhorn
strict-localwith 1 replica. - You are on managed cloud Kubernetes where native CSI disks (EBS, PD, Azure Disk) already give durability.
V1 or V2 Data Engine?¶
A quick decision path based on the v1.12.1 requirements and feature matrix.
flowchart TD
A["New Longhorn volume"] --> B{"Nodes have local NVMe,<br/>kernel 6.7+, spare CPU cores<br/>and 2 GiB huge pages?"}
B -->|"No"| V1["Use V1 (default)"]
B -->|"Yes"| C{"Need strict-local locality,<br/>backing images or<br/>live engine upgrade?"}
C -->|"Yes"| V1
C -->|"No"| D{"Latency or IOPS-sensitive<br/>workload (databases)?"}
D -->|"No"| V1
D -->|"Yes"| V2["Use V2 (SPDK, GA in v1.12)<br/>dataEngine: v2"]
Key Features¶
| Feature | Detail |
|---|---|
| Synchronous replication | Default 3 replicas (per StorageClass numberOfReplicas). Node, zone and disk anti-affinity |
| Incremental snapshots | Up to 250 per volume by default (254 hard limit). Crash-consistent. Optional filesystem freeze |
| Backups | Incremental 2 MiB (or 16 MiB) compressed blocks. S3, NFS, SMB/CIFS, Azure Blob. Several targets |
| Disaster recovery | Standby (DR) volumes in a remote cluster, activated on failover |
| System backup / restore | Backs up Longhorn resources plus volume backups to the default target |
| RWX volumes | NFSv4.1 (NFS-Ganesha) share-manager pod per volume. Fast failover (experimental) |
| Encryption | LUKS2 through dm-crypt, global or per-volume keys, encrypted backups |
| Non-disruptive upgrades | Live engine upgrade for V1 volumes |
| V2 Data Engine | SPDK. NVMe-TCP frontend (UBLK experimental). GA in v1.12.0. IPv6, QoS, fast linked clones |
| Replica rebuild | Delta and fast rebuild. Parallel (scale) rebuild for V1 since v1.11 |
| Networking | Dedicated storage network (Multus). IPv6 and dual-stack (v1.12). Internal NetworkPolicies and gRPC mTLS (v1.12.1) |
| Storage sharding | Erasure-coded V2 layout, experimental in v1.12.1 |
Recent Developments (2025-2026)¶
- v1.12.1 (2026-08-14): New V2 linked-clone design (legacy clones deprecated), experimental erasure-coded storage sharding, V2 CPU allocation through the Kubernetes CPU Manager, internal NetworkPolicies on by default, full instance-manager gRPC mTLS.
- v1.12.0 (2026-06-02): V2 Data Engine GA. IPv6 for V2 volumes, dual-stack clusters,
csi-allowed-topology-keys/strictTopology, default V2 CPU mask raised to 2 cores, V2 backing images removed (use CDI), 16 MiB LUKS2 metadata reservation. - v1.11.0 (2026-01-29): V2 reached Technical Preview. Scale (parallel) replica rebuilding, balance-aware scheduling, S.M.A.R.T. disk health monitoring, RWOP,
allowedTopologies. - v1.10.0 (2025-09-25): V2 interrupt mode and hugepage-free mode, unified V1/V2 settings, single-stack IPv6,
CSIStorageCapacity, configurable backup block size,longhorn.io/v1beta1API removed. - KubeCon EU 2026: SUSE presented Longhorn V2 as its high-performance engine.
Version-by-version facts are in Reference: Release and Support Matrix.
Ecosystem and Commercial Context¶
- SUSE / Rancher: Longhorn is offered as SUSE Storage inside Rancher Prime subscriptions, with its own SUSE support matrix per Longhorn minor. It installs from the Rancher Apps catalog.
- Harvester (SUSE HCI) uses Longhorn, including the V2 engine and migratable RWX volumes, for VM disks.
- Backups commonly target MinIO or other S3-compatible stores. Monitoring uses the Prometheus metrics exposed by
longhorn-manager, commonly visualised in Grafana.
Topic Map¶
- How-to Guides: install, upgrade, StorageClasses, RWX, V2, encryption, backups, system backup, DR, tuning, troubleshooting
- Reference: release matrix, requirements, V1 vs V2 features, setting defaults, backup URL formats, ports, CRDs, benchmarks
- Explanation: architecture, data engines, replication and rebuild, snapshots, CSI flow, RWX, backup/DR internals, security model
Related Topics¶
- Storage Comparison: Ceph vs MinIO vs Longhorn
- Storage domain index
- Ceph: unified block/object/file alternative (Rook on Kubernetes)
- MinIO: S3-compatible backup target
- Kubernetes: CSI and persistent volume model
- PostgreSQL: typical stateful workload on Longhorn volumes
- Cilium and Calico: CNIs validated with Longhorn's internal NetworkPolicies
- Argo CD and Flux: supported GitOps install methods
Sources¶
- Longhorn documentation (latest)
- Architecture and Concepts
- v1.12.1 Important Notes and v1.12.0 Important Notes
- GitHub release v1.12.1 and v1.12.0
- Longhorn repository README (release table, EOL policy)
- Longhorn Helm chart index (release dates)
- What's New in Longhorn 1.11 and What's New in Longhorn 1.10
- Data Engine Comparison
- CNCF project page: Longhorn
- CNCF: Longhorn joins the CNCF Incubator (2021-11-04)
- SUSE Storage in Rancher Manager docs
- SUSE: Longhorn V2 at KubeCon EU 2026
- Harvester: Longhorn V2 Data Engine
Questions¶
Answered¶
- Can Longhorn back up to S3? Yes. S3-compatible stores, NFS, SMB/CIFS and Azure Blob are all supported, and v1.8+ allows several targets. See Backups and Disaster Recovery.
- How does the Longhorn Engine differ from a traditional storage controller? There is one engine per volume, not one per array, so an engine crash affects one volume. Engines live inside the per-node instance-manager pod, which is the wider failure domain. See Instance Manager.
- What is the difference between the V1 and V2 Data Engines? V1 uses iSCSI and Linux processes over sparse files. V2 uses SPDK (RAID bdev engine, lvol replicas on raw block disks) with an NVMe-TCP frontend. V2 is GA since v1.12.0 but still lacks a few V1 features. See Longhorn Engine: V1 and V2.
- How does Longhorn handle replica failure? The failed replica is marked ERR. The manager adds a blank write-only replica after a system snapshot, syncs history in the background, then switches it to read-write. See Replica Rebuilding.
- Does Longhorn support volume encryption? Yes: LUKS2 through dm-crypt, with global or per-PVC Kubernetes Secrets, for Filesystem and Block modes. Backups stay encrypted. See How-to: Set Up Volume Encryption.
- How do Longhorn backups work? A backup flattens a snapshot into compressed, checksum-addressed 2 MiB blocks. Later backups upload only changed blocks, and blocks are shared across backups of the same volume.
- What is a Disaster Recovery (DR) volume? A standby volume in another cluster that keeps restoring the latest backups and is activated on failover. See Disaster Recovery Volumes.
- How does Longhorn provide ReadWriteMany (RWX) access? Through a per-volume share-manager pod running NFS-Ganesha (NFSv4.1) in front of a Longhorn block volume. See Share Manager.
- What is the read index and why does it matter? A 1-byte-per-4 KB in-memory map (V1) of which snapshot layer holds each block. It costs about 256 MB per replica per TB and caps a volume at 254 snapshots.
- Is Longhorn crash-consistent or application-consistent? Crash-consistent. Application consistency needs app-level quiescing, or the optional filesystem-freeze setting.
- When will the V2 Data Engine reach GA? It did in v1.12.0 (2026-06-02), after Technical Preview in v1.11.0.
Open¶
- What are the practical scaling limits beyond ~10 workers and ~1,500 volumes? The official v1.6 scale report stops at 9 workers. No upstream report covers 100+ node clusters or V2 at scale. The v1.12.1 notes flag growing V2 attach latency (issue #13241).
- How does Longhorn V2 compare with Ceph RBD for PostgreSQL/MySQL on the same hardware? Upstream publishes V1/V2 kbench numbers, but a controlled database benchmark against Ceph RBD has not been verified.
- When will V2 support live engine upgrade? Upstream says it is planned for the 1.12 to 1.13 upgrade. No v1.13.0 GA date is published;
longhorn-managerreachedv1.13.0-rc4on 2026-09-23. - Will storage sharding (erasure coding) become production-ready, and will it gain backup/DR support? It is experimental in v1.12.1.