Skip to content

Ceph

Summary

Ceph is an open-source (LGPL), software-defined storage system that serves block (RBD), object (RGW, S3/Swift), and file (CephFS) storage from one self-healing RADOS cluster, plus NVMe/TCP, SMB, and NFS gateways. CRUSH-based placement removes central lookup tables, so it scales from three nodes to more than 10,000 OSDs. The current stable release is Tentacle 20.2.4 (2026-08-19), a security release that also changed CephX key handling. Squid 19.2 is still supported until its 2026-10-31 target EOL, and Umbrella (v21) is at the release-candidate stage.

Key Facts

Attribute Detail
Latest Version 20.2.4 "Tentacle" (2026-08-19)
Other supported series 19.2.6 "Squid" (2026-08-19), target EOL 2026-10-31
Next release v21 "Umbrella" (21.1.x release candidates, 2026-09; no GA date published as of 2026-09-27)
Release cadence One named stable release per year; two active series at a time
Repository github.com/ceph/ceph
Stars ~14k+ (recorded by 2026-08)
Language C++, Python
License LGPL-2.1 or LGPL-3.0 (core)
Governance Ceph Foundation (Linux Foundation directed fund) for funding; Ceph Steering Committee and Executive Council for technical decisions
Major contributors IBM (formerly the Red Hat Ceph team), CLYSO, Canonical, and others
Install paths cephadm (hosts), Rook v1.20.7 (Kubernetes), pveceph (Proxmox VE)

August 2026 security release

Tentacle 20.2.4 and Squid 19.2.6 fix four high-severity CVEs, including a CephX authentication bypass (CVE-2025-30156) and a monitor config-key store leak (CVE-2026-50152). Upgrading is only half the job: every CephX key must then be rotated to the new aes256k type, and secrets in the config-key store should be rotated. Reef and older releases get no fix. See How-to Guides.

Architecture at a Glance

The compact diagram shows how clients reach RADOS; the full component view is in Explanation.

flowchart LR
    subgraph Access["Access layer"]
        RBD["librbd / krbd"]
        RGW["radosgw (S3/Swift)"]
        CFS["CephFS client"]
        GW["NVMe-oF / SMB / NFS gateways"]
    end
    subgraph Core["RADOS"]
        MON["ceph-mon quorum"]
        MGR["ceph-mgr"]
        MDS["ceph-mds"]
        OSD["ceph-osd + BlueStore"]
    end
    RBD --> OSD
    RGW --> OSD
    CFS --> MDS
    CFS --> OSD
    GW --> OSD
    MON -.->|"cluster maps"| RBD
    MON -.->|"cluster maps"| OSD
    MGR --> MON

Evaluation

Pros Cons
Unified block, object, and file from one cluster Complex to design, deploy, and operate
CRUSH: no single lookup service, linear scaling High RAM/CPU needs (default 4 GiB per OSD memory target)
Self-healing, automatic rebalancing Needs tuning and fast networks for good latency
Industry standard for OpenStack; Rook for Kubernetes Small-write EC was slow before Tentacle's FastEC
Exabyte-scale deployments; CERN tested >10,000 OSDs Recovery and backfill can saturate the network
cephadm and Rook give automated, rolling upgrades 2026 CephX migration adds operational work and kernel requirements
FastEC (Tentacle) makes EC viable for more block/file workloads Crimson/SeaStore still tech preview

When Ceph Fits

  • You need more than one storage type (VM disks plus S3 plus shared file) and want one system to operate.
  • You run OpenStack, Proxmox VE hyperconverged clusters, or large Kubernetes estates (via Rook).
  • You expect to scale out past a few hundred TB or need rack/datacenter-level failure domains.

Consider alternatives when you need only Kubernetes block storage on a small cluster (Longhorn is simpler) or only S3 (a dedicated object store may be easier; see MinIO and its licensing caveats).

What's New (2025-2026)

Release Highlights
Tentacle 20.2.0 (2025-11-18) FastEC (partial reads/writes, less padding) opt-in per pool; ISA-L as default EC plugin for new clusters; SMB manager module with AD and CTDB; SeaStore tech preview with Crimson; faster BlueStore WAL and better compression; mgmt-gateway + oauth2-proxy SSO; certmgr; RBD instant import from other clusters or NBD sources; RGW GetObjectAttributes and less-blocking resharding; tenant IAM deprecated in favour of accounts
Tentacle 20.2.1-20.2.3 (2026-04 to 2026-08) NVMe-oF fast-failover overhaul and new nvmeof mgr module; RBD transient exclusive locks; Rocky Linux 10 packages; RGW TLS 1.3 cipher control and Kafka mTLS
20.2.4 / 19.2.6 (2026-08-19) Four CVE fixes and the aes256k CephX key type
Umbrella v21 (RC) Further FastEC work, RGW dedup (tech preview), bucket logging, cloud restore, Crimson/SeaStore EC (draft RC notes; subject to change)

Storage Interfaces

Interface Protocol Use Case
RBD Block (librbd, krbd, Ceph-CSI) VM disks, Kubernetes PVs, databases
RGW Object (S3/Swift API) Backups, media, data lakes
CephFS File (POSIX) Shared filesystems, HPC, RWX volumes
NVMe-oF gateway NVMe/TCP Block for VMware and hosts without Ceph clients
SMB / NFS gateways SMB2/3, NFSv4 on CephFS Windows shares, legacy NFS clients

Topic Map

  • Explanation: RADOS, CRUSH, FastEC, BlueStore, Crimson/SeaStore, data paths, orchestration models, security model, governance
  • How-to Guides: cephadm and Rook deployment, upgrades, CephX aes256k key rotation, pools, CRUSH, RBD/RGW/CephFS, troubleshooting
  • Reference: release/support matrix, feature table, config defaults, ports, CephX capabilities, CVEs, Rook compatibility, hardening checklist

Sources

Questions

Answered

  • When to use EC vs replication? Replication (3x, 200% extra raw capacity) for hot, small-write, latency-sensitive data such as VM disks and databases. EC (for example 4+2, 50% extra) for capacity-oriented data such as RGW archives and backups. Tentacle's FastEC with a 16 KiB+ stripe unit makes EC reasonable for more RBD and CephFS workloads; benchmark first. See Explanation.
  • How many monitors does a production cluster need? At least 3; use 5 for larger or multi-rack clusters. Always an odd number, because Paxos quorum needs a strict majority.
  • What is the minimum replication factor for data safety? size=2 is technically possible, but use size=3, min_size=2. The upstream docs warn that two-copy pools eventually lose data through overlapping failures.
  • How does CRUSH avoid a central lookup table? Clients and OSDs compute placement from the same CRUSH map: hash the object name to a PG, then map the PG to OSDs following the failure-domain hierarchy. See Explanation.
  • What storage backend does the OSD use? BlueStore, the default since Luminous and the only supported backend since Reef (FileStore removed). It writes data to raw devices and keeps metadata in RocksDB on BlueFS.
  • How does Ceph handle OSD failures? Peers report missed heartbeats, the monitors mark the OSD down, and after mon_osd_down_out_interval (10 minutes by default) mark it out, which remaps its PGs and starts backfill. See the state diagram in Explanation.
  • Can Ceph encrypt data at rest? Yes, per-OSD dm-crypt/LUKS via ceph-volume, with keys held in the monitors' config-key store. BlueStore adds checksums, not encryption.
  • What does ceph-mgr do? Hosts the Dashboard, Prometheus exporter, PG autoscaler, balancer, and orchestrator (cephadm/Rook), NFS, SMB, and NVMe-oF modules. Run one active plus at least one standby.
  • How does CephFS MDS scaling work? Multiple active ranks split the directory tree into subtrees and can shard hot directories; standbys provide failover (for example 3 active + 1 standby).
  • What is Crimson, and is it production-ready? A Seastar-based rewrite of the OSD; first tech preview in Squid, with SeaStore added as a tech preview in Tentacle. As of 2026-09 the docs still call it not suitable for production.
  • cephadm or Rook? Both are upstream-recommended. Use Rook when Ceph runs inside Kubernetes (or to connect Kubernetes to an external cluster); use cephadm for dedicated hosts. See the decision flowchart in Explanation.
  • What is Ceph's next release? Umbrella (v21), in release candidates as of 2026-09; Rook v1.20 will need allowUnsupported: true to run it.

Open

  • How does SeaStore compare to BlueStore on NVMe? It needs published benchmarks once SeaStore leaves tech preview.
  • What are FastEC's measured gains? Upstream describes the mechanism and gives tuning guidance (16 KiB+ stripe unit), but a reproducible benchmark against classic EC for RBD 4K random writes has not been captured here.
  • What is the PG autoscaler's overhead on 1000+ OSD clusters? Rebalancing impact under production load still needs validation, especially with main's higher mon_target_pg_per_osd default (200 vs 100 in Tentacle).
  • Exact Squid EOL? Upstream releases.yml targets 2026-10-31. Confirm when the final Squid release ships.
  • When will Umbrella v21 reach GA? Upstream has not published a date (releases.yml has no Umbrella entry as of 2026-09-27); the RC notes are drafts.