Ceph¶
Summary
Ceph is an open-source (LGPL), software-defined storage system that serves block (RBD), object (RGW, S3/Swift), and file (CephFS) storage from one self-healing RADOS cluster, plus NVMe/TCP, SMB, and NFS gateways. CRUSH-based placement removes central lookup tables, so it scales from three nodes to more than 10,000 OSDs. The current stable release is Tentacle 20.2.4 (2026-08-19), a security release that also changed CephX key handling. Squid 19.2 is still supported until its 2026-10-31 target EOL, and Umbrella (v21) is at the release-candidate stage.
Key Facts¶
| Attribute | Detail |
|---|---|
| Latest Version | 20.2.4 "Tentacle" (2026-08-19) |
| Other supported series | 19.2.6 "Squid" (2026-08-19), target EOL 2026-10-31 |
| Next release | v21 "Umbrella" (21.1.x release candidates, 2026-09; no GA date published as of 2026-09-27) |
| Release cadence | One named stable release per year; two active series at a time |
| Repository | github.com/ceph/ceph |
| Stars | ~14k+ (recorded by 2026-08) |
| Language | C++, Python |
| License | LGPL-2.1 or LGPL-3.0 (core) |
| Governance | Ceph Foundation (Linux Foundation directed fund) for funding; Ceph Steering Committee and Executive Council for technical decisions |
| Major contributors | IBM (formerly the Red Hat Ceph team), CLYSO, Canonical, and others |
| Install paths | cephadm (hosts), Rook v1.20.7 (Kubernetes), pveceph (Proxmox VE) |
August 2026 security release
Tentacle 20.2.4 and Squid 19.2.6 fix four high-severity CVEs, including a CephX authentication bypass (CVE-2025-30156) and a monitor config-key store leak (CVE-2026-50152). Upgrading is only half the job: every CephX key must then be rotated to the new aes256k type, and secrets in the config-key store should be rotated. Reef and older releases get no fix. See How-to Guides.
Architecture at a Glance¶
The compact diagram shows how clients reach RADOS; the full component view is in Explanation.
flowchart LR
subgraph Access["Access layer"]
RBD["librbd / krbd"]
RGW["radosgw (S3/Swift)"]
CFS["CephFS client"]
GW["NVMe-oF / SMB / NFS gateways"]
end
subgraph Core["RADOS"]
MON["ceph-mon quorum"]
MGR["ceph-mgr"]
MDS["ceph-mds"]
OSD["ceph-osd + BlueStore"]
end
RBD --> OSD
RGW --> OSD
CFS --> MDS
CFS --> OSD
GW --> OSD
MON -.->|"cluster maps"| RBD
MON -.->|"cluster maps"| OSD
MGR --> MON
Evaluation¶
| Pros | Cons |
|---|---|
| Unified block, object, and file from one cluster | Complex to design, deploy, and operate |
| CRUSH: no single lookup service, linear scaling | High RAM/CPU needs (default 4 GiB per OSD memory target) |
| Self-healing, automatic rebalancing | Needs tuning and fast networks for good latency |
| Industry standard for OpenStack; Rook for Kubernetes | Small-write EC was slow before Tentacle's FastEC |
| Exabyte-scale deployments; CERN tested >10,000 OSDs | Recovery and backfill can saturate the network |
| cephadm and Rook give automated, rolling upgrades | 2026 CephX migration adds operational work and kernel requirements |
| FastEC (Tentacle) makes EC viable for more block/file workloads | Crimson/SeaStore still tech preview |
When Ceph Fits¶
- You need more than one storage type (VM disks plus S3 plus shared file) and want one system to operate.
- You run OpenStack, Proxmox VE hyperconverged clusters, or large Kubernetes estates (via Rook).
- You expect to scale out past a few hundred TB or need rack/datacenter-level failure domains.
Consider alternatives when you need only Kubernetes block storage on a small cluster (Longhorn is simpler) or only S3 (a dedicated object store may be easier; see MinIO and its licensing caveats).
What's New (2025-2026)¶
| Release | Highlights |
|---|---|
| Tentacle 20.2.0 (2025-11-18) | FastEC (partial reads/writes, less padding) opt-in per pool; ISA-L as default EC plugin for new clusters; SMB manager module with AD and CTDB; SeaStore tech preview with Crimson; faster BlueStore WAL and better compression; mgmt-gateway + oauth2-proxy SSO; certmgr; RBD instant import from other clusters or NBD sources; RGW GetObjectAttributes and less-blocking resharding; tenant IAM deprecated in favour of accounts |
| Tentacle 20.2.1-20.2.3 (2026-04 to 2026-08) | NVMe-oF fast-failover overhaul and new nvmeof mgr module; RBD transient exclusive locks; Rocky Linux 10 packages; RGW TLS 1.3 cipher control and Kafka mTLS |
| 20.2.4 / 19.2.6 (2026-08-19) | Four CVE fixes and the aes256k CephX key type |
| Umbrella v21 (RC) | Further FastEC work, RGW dedup (tech preview), bucket logging, cloud restore, Crimson/SeaStore EC (draft RC notes; subject to change) |
Storage Interfaces¶
| Interface | Protocol | Use Case |
|---|---|---|
| RBD | Block (librbd, krbd, Ceph-CSI) | VM disks, Kubernetes PVs, databases |
| RGW | Object (S3/Swift API) | Backups, media, data lakes |
| CephFS | File (POSIX) | Shared filesystems, HPC, RWX volumes |
| NVMe-oF gateway | NVMe/TCP | Block for VMware and hosts without Ceph clients |
| SMB / NFS gateways | SMB2/3, NFSv4 on CephFS | Windows shares, legacy NFS clients |
Topic Map¶
- Explanation: RADOS, CRUSH, FastEC, BlueStore, Crimson/SeaStore, data paths, orchestration models, security model, governance
- How-to Guides: cephadm and Rook deployment, upgrades, CephX
aes256kkey rotation, pools, CRUSH, RBD/RGW/CephFS, troubleshooting - Reference: release/support matrix, feature table, config defaults, ports, CephX capabilities, CVEs, Rook compatibility, hardening checklist
Related Topics¶
- Storage Comparison: Ceph vs MinIO vs Longhorn
- MinIO Alternatives: Ceph RGW as a MinIO replacement
- Longhorn: Kubernetes-native block storage
- MinIO: S3-compatible object storage
- Proxmox VE: hyperconverged Ceph via
pveceph; PVE 9.2 defaults new installs to Tentacle - Kubernetes: Ceph-CSI and Rook for persistent volumes
Sources¶
- Ceph documentation
- Ceph releases index and
releases.yml(dates, EOL targets) - Tentacle release notes
- Squid release notes
- v20.2.0 Tentacle released (Ceph blog)
- Squid v19.2.6 and Tentacle v20.2.4 security release (Ceph blog)
- Ceph security advisories
- Architecture
- BlueStore configuration reference
- Crimson (tech preview)
- Installing Ceph (recommended methods)
- Ceph Foundation and Governance
- Rook documentation and Rook v1.20 release blog
- Proxmox forum: CephX key migration and Squid EOL
- GitHub: ceph/ceph
Questions¶
Answered¶
- When to use EC vs replication? Replication (3x, 200% extra raw capacity) for hot, small-write, latency-sensitive data such as VM disks and databases. EC (for example 4+2, 50% extra) for capacity-oriented data such as RGW archives and backups. Tentacle's FastEC with a 16 KiB+ stripe unit makes EC reasonable for more RBD and CephFS workloads; benchmark first. See Explanation.
- How many monitors does a production cluster need? At least 3; use 5 for larger or multi-rack clusters. Always an odd number, because Paxos quorum needs a strict majority.
- What is the minimum replication factor for data safety?
size=2is technically possible, but usesize=3,min_size=2. The upstream docs warn that two-copy pools eventually lose data through overlapping failures. - How does CRUSH avoid a central lookup table? Clients and OSDs compute placement from the same CRUSH map: hash the object name to a PG, then map the PG to OSDs following the failure-domain hierarchy. See Explanation.
- What storage backend does the OSD use? BlueStore, the default since Luminous and the only supported backend since Reef (FileStore removed). It writes data to raw devices and keeps metadata in RocksDB on BlueFS.
- How does Ceph handle OSD failures? Peers report missed heartbeats, the monitors mark the OSD
down, and aftermon_osd_down_out_interval(10 minutes by default) mark itout, which remaps its PGs and starts backfill. See the state diagram in Explanation. - Can Ceph encrypt data at rest? Yes, per-OSD dm-crypt/LUKS via
ceph-volume, with keys held in the monitors' config-key store. BlueStore adds checksums, not encryption. - What does ceph-mgr do? Hosts the Dashboard, Prometheus exporter, PG autoscaler, balancer, and orchestrator (cephadm/Rook), NFS, SMB, and NVMe-oF modules. Run one active plus at least one standby.
- How does CephFS MDS scaling work? Multiple active ranks split the directory tree into subtrees and can shard hot directories; standbys provide failover (for example 3 active + 1 standby).
- What is Crimson, and is it production-ready? A Seastar-based rewrite of the OSD; first tech preview in Squid, with SeaStore added as a tech preview in Tentacle. As of 2026-09 the docs still call it not suitable for production.
- cephadm or Rook? Both are upstream-recommended. Use Rook when Ceph runs inside Kubernetes (or to connect Kubernetes to an external cluster); use cephadm for dedicated hosts. See the decision flowchart in Explanation.
- What is Ceph's next release? Umbrella (v21), in release candidates as of 2026-09; Rook v1.20 will need
allowUnsupported: trueto run it.
Open¶
- How does SeaStore compare to BlueStore on NVMe? It needs published benchmarks once SeaStore leaves tech preview.
- What are FastEC's measured gains? Upstream describes the mechanism and gives tuning guidance (16 KiB+ stripe unit), but a reproducible benchmark against classic EC for RBD 4K random writes has not been captured here.
- What is the PG autoscaler's overhead on 1000+ OSD clusters? Rebalancing impact under production load still needs validation, especially with main's higher
mon_target_pg_per_osddefault (200 vs 100 in Tentacle). - Exact Squid EOL? Upstream
releases.ymltargets 2026-10-31. Confirm when the final Squid release ships. - When will Umbrella v21 reach GA? Upstream has not published a date (
releases.ymlhas no Umbrella entry as of 2026-09-27); the RC notes are drafts.