Storage¶
Summary
Distributed storage for Kubernetes, virtualization and cloud-native platforms: unified software-defined storage (Ceph), Kubernetes-native replicated block storage (Longhorn), and S3-compatible object storage (MinIO, whose community edition is archived). The 2026 headlines are Ceph's August security release with its mandatory CephX key rotation, Longhorn's SPDK-based V2 Data Engine reaching GA in v1.12, and the end of the MinIO community edition, which pushed many teams to AIStor, the Silo fork or other open-source S3 stores. Versions and licenses below match the topic pages as of 2026-09-25.
Domain Map¶
The map shows how the three systems relate on Kubernetes and as S3 targets. Solid arrows are deployment or data paths; dotted arrows are replacement options.
flowchart LR
subgraph K8S["Kubernetes (CSI)"]
LHCSI["driver.longhorn.io"]
CEPHCSI["Ceph-CSI<br/>(RBD, CephFS)"]
ROOK["Rook v1.20<br/>operator"]
end
LH["Longhorn 1.12.1<br/>block: V1 iSCSI / V2 SPDK"]
CEPH["Ceph 20.2.4 Tentacle<br/>RBD + RGW + CephFS"]
MINIO["MinIO<br/>community archived / AIStor"]
ALT["Open-source S3 stores<br/>Silo, RustFS, SeaweedFS, Garage"]
LHCSI --> LH
CEPHCSI --> CEPH
ROOK -->|"deploys"| CEPH
LH -->|"backups to S3 target"| MINIO
LH -->|"backups to S3 target"| CEPH
CEPH -.->|"RGW replaces"| MINIO
ALT -.->|"replace"| MINIO
Topics¶
| Topic | Type | Summary | Latest version (2026-09-25) | License |
|---|---|---|---|---|
| Ceph | Unified block, object, file | Block (RBD), object (RGW, S3/Swift) and file (CephFS) from one self-healing RADOS cluster placed by CRUSH, plus NVMe/TCP, SMB and NFS gateways. Deployed with cephadm, Rook or Proxmox pveceph |
20.2.4 "Tentacle" (2026-08-19). Squid 19.2.6 until its 2026-10-31 target EOL | LGPL-2.1 or LGPL-3.0 |
| Longhorn | Kubernetes block | Per-volume engine with synchronous replicas, incremental backups to S3/NFS/SMB/Azure, DR volumes, RWX through an NFS share manager. Engines: iSCSI (V1, default) or SPDK with NVMe-TCP (V2, GA since v1.12.0) | v1.12.1 (2026-08-14) | Apache-2.0 (CNCF Incubating) |
| MinIO | S3 object | S3-compatible object store with erasure coding, bitrot checks, IAM/STS and site replication. Community edition archived in 2026; development continues only in the commercial AIStor | Community: RELEASE.2025-10-15T17-29-55Z (last). AIStor: RELEASE.2026-08-07T18-34-35Z (newest seen) |
AGPLv3 (community, archived). AIStor: proprietary (Free tier) or commercial |
Security and support notices (2026)
- Ceph: 20.2.4 and 19.2.6 (2026-08-19) fix four high-severity CVEs, including a CephX authentication bypass. After upgrading, rotate every CephX key to
aes256k. Reef and older get no fix. See Ceph How-to. - MinIO: the community repository was archived in February 2026 (re-archived on 2026-04-25, as reported) and the Docker Hub images were deleted on 2026-09-11. CVE-2026-39414 is fixed only in AIStor. See MinIO.
Comparisons¶
| Comparison | Scope |
|---|---|
| Storage Comparison | Ceph vs Longhorn vs MinIO: interfaces, architecture, minimum footprint, use cases, operations, with a decision flowchart |
| MinIO Alternatives | PGSTY Silo vs RustFS vs SeaweedFS vs Garage vs Ceph RGW: license, status, migration path, with a decision flowchart |
When to Use Which¶
| Need | Reach for |
|---|---|
| Block, object and file from one system; OpenStack or Proxmox VE hyperconverged storage | Ceph |
| Replicated Kubernetes volumes on a small or medium cluster without a storage team | Longhorn (V1 engine) |
| Low-latency Kubernetes volumes on local NVMe with spare CPU cores and huge pages | Longhorn V2 or Ceph RBD; benchmark both |
| Large Kubernetes estates or rack-level failure domains | Ceph via Rook |
| Shared POSIX filesystem | CephFS (Longhorn RWX is NFS in front of one volume) |
| Open-source S3 object storage | Ceph RGW, or see MinIO Alternatives |
| S3 object storage with vendor support | MinIO AIStor (Free tier is single node only) |
| Existing MinIO cluster, no budget | PGSTY Silo fork (same on-disk format), or migrate with mc mirror |
Details and caveats are in the Storage Comparison decision flowchart.
Landscape¶
The Kubernetes storage ecosystem matured around the Container Storage Interface (CSI) specification. CSI decoupled storage provisioning from the kubelet and enabled a large catalogue of drivers covering cloud-provider volumes, SAN/NAS appliances, and software-defined storage systems. S3-compatible object storage is now a common persistence layer for distributed systems. It holds unstructured data and also backs observability platforms (Loki, Tempo, Mimir, OpenObserve), data lakes (Iceberg, Delta Lake), and AI/ML training pipelines.
Software-defined storage (Ceph, Longhorn, and others such as OpenEBS, which has no topic page here) abstracts away hardware heterogeneity. Organizations can then build storage tiers on commodity servers with NVMe, SSD, and HDD pools managed by placement policies. Erasure coding is the storage-efficient alternative to full replication for large-scale object stores. MinIO protects objects with Reed-Solomon erasure coding, and Ceph supports EC pools for RGW, CephFS and, since Tentacle's FastEC, more RBD workloads.
Storage as a Platform
Ceph evolved from a POSIX filesystem into a unified block/object/file platform that can be managed declaratively through the Rook Kubernetes operator (a CNCF graduated project). Rook translates CephCluster, CephBlockPool, and CephObjectStore CRDs into the underlying Ceph configuration; since Rook v1.20 the Ceph-CSI operator is the only supported way to configure the CSI drivers. This enables GitOps-driven storage management alongside application workloads.
Data protection and disaster recovery are increasingly built into the storage layer. Longhorn provides scheduled snapshots and backups to S3, NFS, SMB or Azure Blob targets plus DR standby volumes; Ceph supports RBD mirroring and RGW multisite for cross-cluster DR; MinIO (AIStor and forks) offers site replication for active-active multi-site deployments. NVMe over TCP is reaching both systems in this domain: Ceph exports RBD images through its NVMe-oF gateway, and Longhorn's V2 engine uses an NVMe-TCP frontend between the engine and the node.
Key Concepts¶
CSI (Container Storage Interface)¶
CSI is a standardized gRPC-based interface between container orchestrators (Kubernetes) and storage providers. Storage vendors can implement a single plugin that works across any CSI-compatible orchestrator.
CSI defines three core services, plus an optional GroupController service for volume group snapshots:
- Identity: Capability reporting and plugin health checks
- Controller: Volume create/delete, snapshot create/delete, volume expand, and topology awareness
- Node: Mount/unmount, stage/unstage, and node-local operations
Kubernetes exposes CSI through StorageClass, PersistentVolume, PersistentVolumeClaim, and VolumeSnapshot objects. The controller plugin usually runs as a Deployment alongside Kubernetes CSI sidecar containers (provisioner, attacher, snapshotter, resizer), and the node plugin runs as a DaemonSet.
Block vs Object vs File¶
Storage Modalities
- Block storage (RBD, Longhorn, EBS, Persistent Disks): Raw, fixed-size blocks exposed as a disk device. Ideal for databases and stateful workloads that need low-latency access. Mounted by a single pod (ReadWriteOnce) in most implementations.
- Object storage (S3, RGW, MinIO): Flat namespace of key-value blobs accessed via HTTP APIs (PUT/GET/DELETE). Scales to billions of objects with built-in metadata. No POSIX semantics, so not mountable as a filesystem without FUSE adapters (s3fs, goofys). Dominant for analytics, backups, and observability data.
- File storage (CephFS, NFS, EFS): Hierarchical filesystem accessible by multiple clients simultaneously (ReadWriteMany). Necessary for shared workloads like CMS platforms, legacy applications, and build caches. CephFS provides POSIX compliance at scale using MDS (Metadata Server) for namespace operations. Longhorn RWX volumes put an NFSv4.1 server in front of a single block volume instead.
Erasure Coding¶
A data protection scheme that splits data into k data fragments and m parity fragments. It tolerates up to m simultaneous drive failures while using only (k+m)/k times the raw data size, which is much more storage-efficient than full replication.
- MinIO: default parity is
EC:4for erasure sets of 8 to 16 drives. On a 16-drive set that is 12 data + 4 parity shards, a 1.33x overhead (75% usable). The maximum isEC:N/2(8 + 8 on 16 drives, 2x), and AIStor docs requireEC:3or higher in production. See MinIO Reference. - Ceph: the default EC profile on fresh clusters is
k=2 m=2(2x); a common 4+2 profile gives 1.5x. EC pools suit RGW and CephFS. Replication remains the default for latency-sensitive block (RBD) workloads because classic EC turns small writes into read-modify-write cycles; Tentacle's opt-in FastEC adds partial reads and writes that narrow the gap. See Ceph Reference.
The trade-off is always storage efficiency versus read/write latency and CPU overhead for parity computation.
Replication Factor¶
The number of copies of each data unit maintained across storage nodes for durability and availability. Ceph defaults to 3x replication for pool type "replicated" (size=3, min_size=2); each write is committed to a primary OSD and two replicas before acknowledging the client. Longhorn similarly defaults to 3 replicas with synchronous writes (upstream advises 2 when fewer than 3 storage nodes exist). Higher replication factors increase durability (tolerating more simultaneous node failures) but linearly increase storage consumption and write amplification. The choice between replication and erasure coding depends on the I/O pattern of the workload: replication for latency-sensitive random I/O (databases), erasure coding for throughput-oriented sequential I/O (object storage, backups).
Storage Classes¶
Kubernetes StorageClass objects define tiers of storage with different performance, durability, and cost characteristics. Each StorageClass specifies a CSI provisioner, reclaim policy (Delete or Retain), volume binding mode (Immediate or WaitForFirstConsumer), and provider-specific parameters (IOPS, throughput, filesystem type, encryption). Platform teams often expose a handful of classes, for example fast (NVMe-backed, high IOPS for databases), standard (SSD, general purpose), bulk (HDD or erasure-coded, high capacity for logs/backups), and shared (ReadWriteMany file storage). WaitForFirstConsumer binding makes sure that volumes are provisioned in the same availability zone as the scheduled pod. This prevents cross-AZ latency.
Related¶
- Kubernetes: CSI, StorageClasses and the Rook and Longhorn operators
- Proxmox VE: hyperconverged Ceph via
pveceph - OpenStack: Ceph RBD as the standard block backend
- Grafana LGTM and OpenObserve: object storage as the backend for logs, traces and metrics
- Apache Pulsar: tiered storage to S3-compatible stores
- PostgreSQL: a typical stateful workload on Longhorn or Ceph RBD volumes
Sources¶
- Ceph documentation and Ceph releases
- Squid v19.2.6 and Tentacle v20.2.4 security release (Ceph blog)
- Ceph RBD mirroring and RGW multisite
- Rook documentation
- Longhorn documentation and v1.12.1 Important Notes
- MinIO AIStor product page and AIStor documentation
- minio/minio repository (archived)
- Container Storage Interface specification
Open Questions¶
- As S3-compatible APIs become the universal storage interface for distributed systems, does block storage (RBD, Longhorn) remain necessary beyond stateful database workloads, or will S3-native architectures (Iceberg, Delta Lake) subsume most persistent storage needs?
- How does the operational complexity of Ceph compare to running separate purpose-built systems (Longhorn for block, an S3 store for object)? At what scale does the unified architecture of Ceph justify the operational investment?
- With NVMe-oF (NVMe over Fabrics) enabling remote NVMe access with near-local latency, will software-defined storage on commodity hardware lose its cost advantage to disaggregated NVMe pools managed by dedicated controllers?
- Which open-source S3 store best replaces MinIO on S3 API coverage and performance? The MinIO Alternatives page covers licenses and migration paths, but no controlled benchmark is recorded yet.
- How does Longhorn V2 compare with Ceph RBD for databases on the same NVMe hardware? Neither topic page has a controlled benchmark.