Explanation¶
Scope
How Ceph works and why it is built this way: RADOS and its daemons, CRUSH placement, replication and erasure coding (including Tentacle's FastEC), BlueStore, the Crimson/SeaStore rewrite, data paths, gateways, orchestration models, the security model (including the 2026 CephX redesign), performance behaviour, and governance. Look-up tables live in Reference; tasks live in How-to Guides.
Ceph is a unified distributed storage system that delivers object, block, and file storage from a single RADOS (Reliable Autonomic Distributed Object Store) cluster. Intelligent daemons and the CRUSH algorithm remove centralized lookup bottlenecks, which lets clusters grow from a few nodes to tens of thousands of OSDs on commodity hardware.
Architecture Overview¶
The component diagram below shows the RADOS core, the daemons that manage it, and the gateways and client libraries that expose it as block, object, and file storage.
flowchart TB
subgraph Clients["Clients and protocols"]
KRBD["krbd / librbd<br/>(QEMU, Ceph-CSI)"]
S3C["S3 / Swift clients"]
FSC["CephFS kernel / ceph-fuse"]
NVMEI["NVMe/TCP initiators<br/>(Linux, VMware)"]
SMBC["SMB clients"]
NFSC["NFS clients"]
end
subgraph Gateways["Gateway daemons"]
RGW["radosgw (RGW)<br/>Beast frontend"]
NVMEGW["NVMe-oF gateway<br/>SPDK + bdev_rbd"]
SMBD["Samba + CTDB<br/>(smb mgr module)"]
NFSG["NFS-Ganesha<br/>(nfs mgr module)"]
end
subgraph RADOS["RADOS cluster"]
MON["ceph-mon x3/5<br/>Paxos, cluster maps, CephX"]
MGR["ceph-mgr active+standby<br/>dashboard, prometheus, cephadm/rook"]
MDS["ceph-mds<br/>CephFS metadata"]
OSD["ceph-osd x N<br/>BlueStore on raw devices"]
end
KRBD --> OSD
S3C --> RGW --> OSD
FSC --> MDS
FSC --> OSD
NVMEI --> NVMEGW --> OSD
SMBC --> SMBD --> FSC
NFSC --> NFSG --> MDS
MDS --> OSD
MON -->|"maps + tickets"| OSD
MON -->|"maps + tickets"| Clients
MGR --> MON
Key architectural properties:
- No centralized data gateway: native clients compute placement via CRUSH and talk directly to OSDs.
- Self-healing: OSDs detect failures, peer, and rebuild data from surviving replicas or shards automatically.
- Elastic scaling: adding or removing OSDs moves only the PGs whose mapping changes.
- Commodity hardware: no proprietary hardware; each OSD owns its device (shared-nothing).
RADOS Daemons¶
Ceph Monitor (ceph-mon)¶
Monitors maintain the master copy of the cluster map, a set of sub-maps (monitor, OSD, PG, CRUSH, and MDS maps, listed in Reference). They also act as the CephX authentication authority and hold the config-key store (used for secrets such as dm-crypt keys and, under cephadm, the orchestrator's SSH key).
Monitors reach consensus with Paxos. A quorum needs a strict majority (2 of 3, 3 of 5). Losing quorum blocks cluster-map updates and new client sessions, so production clusters run 3 or 5 monitors across failure domains, always an odd number.
Ceph OSD Daemon (ceph-osd)¶
Each OSD manages one storage device (typically one per physical disk or NVMe namespace) and does the heavy lifting: reads, writes, replication, recovery, backfill, rebalancing, and scrubbing.
- Data storage via BlueStore (the default backend since Luminous).
- Heartbeats to peer OSDs. Peers report a silent OSD to the monitors, which mark it
down. - Peering and recovery when the acting set of a PG changes.
- Scrubbing: light scrubs compare metadata (default interval 1 day); deep scrubs read and checksum all data (default 7 days).
An OSD has two independent states: up/down (daemon running or not) and in/out (participating in data placement or not).
Ceph Manager (ceph-mgr)¶
The manager hosts Python modules: the Dashboard, the Prometheus exporter, the PG autoscaler, the balancer, and orchestrator backends (cephadm, rook), plus the nfs, smb, and (since 20.2.3) nvmeof modules. One active manager is required; standbys take over automatically. Tentacle removed the long-deprecated restful and zabbix modules.
Ceph Metadata Server (ceph-mds)¶
MDS is only needed for CephFS. It caches filesystem metadata (directories, ownership, permissions, timestamps) for fast POSIX operations and journals changes into RADOS.
- Active-standby HA: one active MDS per rank; standbys take over on failure.
- Active-active scaling: multiple active ranks split the directory tree into subtrees and can shard hot directories. A common layout is 3 active + 1 standby.
RADOS Gateway (radosgw)¶
RGW provides S3- and Swift-compatible REST APIs on top of RADOS, with its own user database, IAM, and access control. Squid added User Accounts, which bring self-service AWS-style IAM APIs (users, groups, roles, policies). Tentacle deprecates the older tenant-level IAM APIs in favour of accounts (removal no sooner than the "V" release), adds GetObjectAttributes, lets Object Lock be enabled on existing versioned buckets, and performs most bucket-resharding work before blocking writes.
CRUSH Algorithm and Placement Groups¶
CRUSH (Controlled Replication Under Scalable Hashing) decides where each object lives. Every client and OSD computes placement independently from the same CRUSH map, so there is no central lookup table to scale or lose.
The flowchart shows how an object name becomes a set of OSDs.
flowchart LR
OBJ["Object<br/>(pool id + object name)"]
HASH["rjenkins hash of name<br/>mod pg_num"]
PG["Placement Group<br/>e.g. 4.58"]
RULE["CRUSH rule<br/>(take root, chooseleaf host)"]
OSDS["Acting set<br/>[osd.4 primary, osd.17, osd.29]"]
OBJ --> HASH --> PG --> RULE --> OSDS
- The client hashes the object name and takes it modulo the pool's PG count, yielding a PG number (for example
58). - The pool ID is prepended (pool
4gives PG4.58). - CRUSH maps the PG to an ordered list of OSDs, the acting set. The first OSD is the primary and coordinates writes.
- Because CRUSH is deterministic, any client with the same map computes the same result.
Placement Groups¶
PGs are logical buckets that amortize the cost of tracking per-object placement. The PG autoscaler (on by default for new pools since Octopus) sizes each pool's pg_num from its usage and a per-OSD target (mon_target_pg_per_osd, see Reference).
- Acting set: the OSDs responsible for a PG; the first is primary.
- Up set: the OSDs CRUSH currently maps the PG to. It differs from the acting set temporarily during backfill (
pg_temp). - Peering: the OSDs of a PG agree on the authoritative state of all objects before the PG becomes
active+clean.
Failure Domains and Rebalancing¶
CRUSH uses a hierarchy of bucket types (osd, host, chassis, rack, row, room, datacenter, root). A rule such as chooseleaf firstn 0 type host guarantees that no two replicas of a PG land on the same host. When OSDs join or leave, only PGs whose mapping changes move data; new OSDs receive a proportional share.
OSD Failure Handling¶
The state diagram shows how an OSD moves through up/down and in/out after a failure, and when data movement starts.
stateDiagram-v2
[*] --> UpIn
UpIn --> DownIn: peers report missed heartbeats, mon marks down
DownIn --> UpIn: daemon restarts within grace
DownIn --> DownOut: mon_osd_down_out_interval expires (default 10 min)
DownOut --> UpOut: daemon restarts
UpOut --> UpIn: ceph osd in, or automatic mark-in
UpIn --> UpOut: operator runs ceph osd out
DownOut --> [*]: ceph osd purge
While an OSD is down but still in, its PGs are degraded but no data moves. Once it is marked out, CRUSH remaps its PGs and backfill rebuilds the missing copies elsewhere. Setting noout during maintenance prevents that rebalancing.
Pools, Replication, and Erasure Coding¶
Replicated Pools¶
Pools default to replication (size = 3, min_size = 2). The primary OSD writes locally and forwards to the replicas in parallel, acknowledging the client only after every replica in the acting set has committed. This gives the lowest write latency and simplest recovery.
Erasure-Coded Pools¶
Erasure coding splits each object into k data chunks and m coding chunks stored on k+m OSDs. A k=3, m=2 profile uses 5 OSDs, tolerates 2 failures, and costs 67% extra raw capacity (versus 200% for 3x replication).
The diagram follows one object through a 3+2 profile.
flowchart LR
OBJ["Object NYAN<br/>ABCDEFGHI"]
subgraph EC["Primary OSD encodes (k=3, m=2)"]
D1["D1: ABC"]
D2["D2: DEF"]
D3["D3: GHI"]
C1["C1: parity"]
C2["C2: parity"]
end
OBJ --> D1
OBJ --> D2
OBJ --> D3
OBJ --> C1
OBJ --> C2
D1 --> O1[("osd.5 shard 0")]
D2 --> O2[("osd.1 shard 1")]
D3 --> O3[("osd.2 shard 2")]
C1 --> O4[("osd.3 shard 3")]
C2 --> O5[("osd.4 shard 4")]
Reads need any k chunks. Writes go through the primary, which encodes and distributes shards. RBD and CephFS can place data in an EC pool (with allow_ec_overwrites) while keeping metadata in a replicated pool.
FastEC in Tentacle¶
Classic EC was a poor fit for block and file workloads: every small write became a read-modify-write of a full stripe, and small objects were padded to the stripe width. Tentacle ships a new EC I/O path ("FastEC", enabled per pool with allow_ec_optimizations) that adds partial reads and partial writes and eliminates padding. The upstream docs describe substantial gains for RBD and CephFS and for RGW workloads with many small objects, and little benefit for large sequential RGW objects.
Design consequences worth knowing:
- The flag needs all monitors and OSDs on Tentacle; clients and gateways do not need upgrading.
- It is one-way: once set, it cannot be cleared because new data is stored differently.
- The stripe unit is fixed at pool creation. The docs recommend at least 16 KiB with optimizations, so existing 4 KiB pools get only part of the benefit.
- With optimizations, raising
kcosts little performance, which makes wider profiles (for example 8+3) more viable. Raisingmstill hurts small writes. - 20.2.1 blocks enabling FastEC on non-4K-aligned chunk sizes after bugs and poor performance were found there.
Tentacle also switches the default EC plugin for new clusters from Jerasure (no longer maintained) to Intel ISA-L; upgraded clusters keep their existing default profile. The Umbrella (v21) release candidates extend the optimized path with direct/sync reads, OMAP support, and deep scrub.
Replication vs erasure coding
Before Tentacle, the rule was simple: replicated pools for hot, latency-sensitive data (RBD VM disks, databases) and EC for cold data (RGW archives, backups). FastEC narrows the gap for block and file, but replication still wins on small-write latency and recovery cost. Benchmark your workload before moving VM disks to EC.
BlueStore¶
BlueStore is the default OSD backend since Luminous and the only supported one since Reef (FileStore removed). It writes object data directly to a raw block device, with no intervening filesystem, and stores metadata in an embedded RocksDB running on BlueFS, a minimal filesystem inside BlueStore.
| Component | Role |
|---|---|
| Block (data) device | Raw HDD/SSD/NVMe holding object data |
| RocksDB on BlueFS | Object metadata (onodes), omap, allocation state |
| WAL device (optional) | RocksDB write-ahead log; best on the fastest device |
| DB device (optional) | RocksDB SST files; SSD/NVMe in front of HDD data |
Design points:
- Direct I/O and its own cache: BlueStore bypasses the page cache and autotunes its caches to
osd_memory_target(4 GiB default). - Checksums on every block: CRC32C by default, verified on read and by deep scrub, which detects silent corruption.
- Inline compression (
zlib,zstd,lz4,snappy) per pool or globally; off by default. Tentacle improves compression and adds a new, faster WAL. - RocksDB LZ4 compression has been enabled by default since Squid to save fast-device space.
- Encryption is done below BlueStore with dm-crypt/LUKS on the OSD's logical volumes (see Security Model); BlueStore itself does not encrypt.
Crimson and SeaStore¶
Crimson is a rewrite of ceph-osd on the Seastar framework: a shared-nothing, thread-per-core, run-to-completion design that avoids locks and context switches so an OSD can keep up with fast NVMe and networks. It is meant to be a drop-in replacement for the classic OSD.
- Squid (19.2): first Crimson tech preview, supporting RBD workloads on replicated pools with BlueStore underneath.
- Tentacle (20.2): SeaStore, Crimson's native object store, is deployable alongside Crimson-OSD as a tech preview. SeaStore is a log-structured design aimed at NVMe and "may not be suitable for traditional HDDs"; Crimson keeps BlueStore support for HDDs and slower SSDs.
- Umbrella (v21, in RC): continued Crimson/SeaStore erasure-coding and performance work.
Not for production
As of 2026-09, the upstream docs still label Crimson a tech preview that is not suitable for production. Enabling it requires the enable_experimental_unrecoverable_data_corrupting_features flag, and cephadm SeaStore support is in early stages.
Data Paths¶
Write Path (Replicated Pool)¶
The sequence shows a client write to a size-3 replicated pool; the client gets its acknowledgment only after all three OSDs have committed.
sequenceDiagram
participant C as librados client
participant P as Primary OSD (osd.4)
participant R1 as Replica OSD (osd.17)
participant R2 as Replica OSD (osd.29)
participant BS as BlueStore (each OSD)
C->>C: CRUSH(pool, object) gives PG 4.58, acting set
C->>P: MOSDOp write
par Replicate and commit locally
P->>BS: Commit data + RocksDB metadata (WAL)
P->>R1: MOSDRepOp
P->>R2: MOSDRepOp
end
R1-->>P: commit reply
R2-->>P: commit reply
P-->>C: MOSDOpReply (committed)
Note over C,P: Ack only after every acting-set OSD commits (strong consistency)
Read Path¶
Clients fetch cluster maps from a monitor when they connect and then receive map updates by subscription; they do not ask a monitor for each I/O.
sequenceDiagram
participant C as Client
participant M as ceph-mon
participant P as Primary OSD
participant BS as BlueStore
C->>M: Authenticate (CephX) and subscribe to OSD map
M-->>C: OSD map, CRUSH map, tickets
C->>C: CRUSH(object) gives PG and acting set
C->>P: Read object
P->>BS: Look up onode in RocksDB
BS-->>P: Extent locations and checksums
P->>BS: Read extents from block device, verify CRC
P-->>C: Object data
Reads go to the primary by default. RBD can spread reads to replicas (rbd_read_from_replica_policy = balance or localize). For EC pools the primary gathers k shards (FastEC can read only the shards that hold the requested range) and reconstructs data if shards are missing.
Access Protocols and Gateways¶
| Interface | Daemon/Lib | Access Method | Typical Use Case |
|---|---|---|---|
| RBD (block) | librbd, kernel rbd |
QEMU/libvirt, krbd, Ceph-CSI | VM disks, Kubernetes PVs |
| RGW (object) | radosgw |
S3 / Swift REST | S3-compatible object storage, backups, data lakes |
| CephFS (file) | ceph-mds, libcephfs |
Kernel client / FUSE | Shared POSIX filesystem, HPC, RWX PVs |
| NVMe-oF | ceph-nvmeof gateway (SPDK) |
NVMe/TCP | Block for hosts without Ceph clients (VMware, Windows, bare metal) |
| SMB | Samba + CTDB via smb mgr module |
SMB2/3 | Windows file shares on CephFS, AD-joined |
| NFS | NFS-Ganesha via nfs mgr module |
NFSv4 | NFS exports of CephFS or RGW buckets |
| librados | librados |
Native object API | Custom applications |
NVMe-oF gateway. Each gateway runs an SPDK NVMe-oF target with the bdev_rbd module plus a control daemon, exporting RBD images as NVMe namespaces. Gateways form gateway groups (up to 8 gateways, up to 4 groups) with active/standby ownership per namespace and automatic failover coordinated by a monitor-side NVMe-oF map. Tentacle added multiple namespaces and dashboard support, 20.2.1 overhauled fast failover, and 20.2.3 introduced an nvmeof mgr module that stores gateway state in a .nvmeof pool. It replaces the iSCSI gateway, which has been in maintenance mode since November 2022.
SMB. The Tentacle smb module works like the NFS module: it deploys Samba containers from the samba-container project, joins Active Directory or uses local users, clusters them with CTDB, and places a cephfs-proxy daemon between Samba and CephFS to reduce memory use.
Deployment and Orchestration Models¶
Ceph has two upstream-recommended installers that share the same orchestrator API in ceph-mgr, so ceph orch commands and the Dashboard work with either.
| Aspect | cephadm | Rook |
|---|---|---|
| Target | Bare-metal or VM hosts | Kubernetes clusters |
| Mechanism | mgr module + SSH to hosts; daemons as Podman/Docker containers under systemd | Kubernetes operator reconciling CRDs (CephCluster, CephBlockPool, CephObjectStore, CephFilesystem) |
| Upgrades | ceph orch upgrade start --image ... (staggered, automated) |
Change spec.cephVersion.image; operator rolls daemons |
| Client integration | Native clients, NFS/SMB/NVMe-oF services | Ceph-CSI (via the Ceph-CSI operator since Rook v1.20), COSI |
| Governance | Part of Ceph | CNCF graduated project |
Proxmox VE has its own integration (pveceph) that installs packages from Proxmox's Ceph repositories; PVE 9.2 defaults new installs to Tentacle (see Proxmox VE).
The decision flowchart helps choose a deployment path.
flowchart TD
Q1{"Where will Ceph run?"}
Q1 -->|"Inside a Kubernetes cluster"| ROOK["Rook operator<br/>(Helm or manifests)"]
Q1 -->|"Dedicated hosts or VMs"| Q2{"Hyperconverged with Proxmox VE?"}
Q2 -->|"Yes"| PVE["pveceph<br/>(Proxmox Ceph repos)"]
Q2 -->|"No"| CEPHADM["cephadm<br/>(containers + systemd)"]
ROOK --> Q3{"Need an existing external cluster?"}
Q3 -->|"Yes"| EXT["Rook external mode<br/>connecting to a cephadm cluster"]
Q3 -->|"No"| INT["Rook-managed cluster<br/>(host or PVC-based OSDs)"]
Security Model¶
Ceph layers several controls: CephX for identity, capabilities for authorization, msgr2 for transport integrity and encryption, dm-crypt for data at rest, and RGW's own IAM for S3 users.
CephX Authentication¶
CephX is Ceph's native, Kerberos-like mutual authentication protocol, enabled by default. The sequence shows how a client obtains tickets.
sequenceDiagram
participant C as client.app1
participant M as ceph-mon
participant O as ceph-osd / ceph-mds
C->>M: Auth request with entity name
M-->>C: Session key, encrypted with client secret
C->>C: Decrypt session key (proves key possession)
C->>M: Request tickets for osd, mds, mgr
M-->>C: Service tickets (sealed with rotating service keys)
C->>O: Present ticket, sign messages
O->>O: Verify ticket with rotating service key
Note over M,O: Rotating service keys are shared by mons and daemons, tickets expire (auth_service_ticket_ttl, 1 h)
CephX authenticates connections between clients and daemons; it does not by itself encrypt data in transit or at rest.
The 2026 CephX Key-Type Change¶
In August 2026, Tentacle 20.2.4 and Squid 19.2.6 fixed CVE-2025-30156. CephX used AES-128-CBC with no MAC and a hard-coded IV, so an attacker holding any low-privilege key (or able to observe CephX traffic) could flip bits in encrypted tickets, for example the allow_all field, and forge administrative credentials. The fix introduces a new key type, aes256k (AES256-CTS-HMAC-SHA384-192, per RFC 8009), with a random confounder per encryption and HMAC authentication. It is the first new CephX key type in Ceph's history.
Design consequences:
- Fresh installs use
aes256k. Upgraded clusters keep working withaeskeys but raise newAUTH_INSECURE_*health warnings until every key is rotated. - cephadm automates daemon-key rotation but not client keys; Rook automates rotation for some client keys.
- Kernel clients (krbd, CephFS kernel mount) need Linux 7.0+ or a vendor backport to use
aes256kkeys, which gates client-key rotation. - The same release fixed CVE-2026-50152: any key with
mon allow rcould dump the monitor config-key store, which holds OSD LUKS passphrases and, on cephadm clusters, the cephadm SSH private key. Upgrading closes the hole but does not un-leak secrets, so those secrets need rotating separately.
The migration steps are in How-to Guides.
Authorization¶
Users are named client.<name> and hold a secret key plus capabilities per daemon type (mon, osd, mds, mgr), which can be scoped to pools and namespaces or use predefined profiles such as profile rbd. See Reference for the keyword table.
Encryption at Rest¶
ceph-volume encrypts OSD logical volumes with dm-crypt (LUKS). At OSD creation it sends the monitors a JSON payload holding the OSD's cephx_secret, the dmcrypt_key, and a cephx_lockbox_secret (a legacy name from ceph-disk) used to retrieve the dm-crypt key. At activation the OSD fetches the key from the monitors, opens the LUKS volume, and starts. DB and WAL volumes are encrypted with the same key. Because the key lives in the monitors, a stolen disk alone cannot be decrypted, which is also why CVE-2026-50152 was serious.
Transport Security (msgr2)¶
The Messenger v2 protocol (monitor port 3300) supports two modes:
crc: CRC32C integrity plus CephX authentication; payloads are not encrypted.secure: payloads encrypted with AES-GCM using keys derived from the CephX session.
The defaults allow both (crc secure, in preference order). RGW terminates TLS itself (Beast ssl_port), behind a load balancer, or through the cephadm ingress and certmgr in Tentacle.
RGW Access Control¶
RGW keeps its own users (access key + secret, AWS SigV4) separate from CephX. It supports S3 bucket policies (including ArnEquals/ArnLike conditions since Tentacle), legacy ACLs, Block Public Access (RestrictPublicBuckets added in Tentacle), STS/OIDC, and account-level IAM. The August 2026 releases also fixed RGW STS token forgery (CVE-2026-39944) and a SigV4 verification flaw (CVE-2026-54330).
Rook Security Context¶
On Kubernetes, OSD pods need privileged access for block devices and dm-crypt, so the Rook namespace needs the privileged Pod Security Standard. The operator and Ceph-CSI use separate ServiceAccounts and RBAC, OpenShift gets SCC definitions, and Rook ships no default NetworkPolicies, so administrators should restrict traffic to the Ceph ports (3300, 6789, 6800-7568).
Threat Model Summary¶
| Threat | Mitigation |
|---|---|
| Unauthorized client access | CephX mutual authentication, capability scoping |
| Ticket forgery by a low-privilege key holder | aes256k keys (20.2.4 / 19.2.6), rotate all keys |
| Secret exfiltration from the config-key store | Upgrade (CVE-2026-50152), then rotate LUKS keys and the cephadm SSH key |
| Man-in-the-middle on cluster network | msgr2 secure mode, isolated cluster network |
| Data exposure from stolen disks | dm-crypt OSD encryption with keys held by the monitors |
| Unauthorized S3 access | RGW access keys, bucket policies, IAM accounts, Block Public Access |
| Privilege escalation in Kubernetes | Rook RBAC, Pod Security Standards, dedicated ServiceAccounts |
| Replay attacks | CephX ticket TTL, rotating service keys |
| Silent data corruption | BlueStore checksums, deep scrubbing |
Performance Characteristics¶
Ceph performance depends on the device class, network, replication versus EC choice, and client type far more than on the software version. Some principles hold across deployments:
- Replication write latency is bounded by the slowest OSD in the acting set plus two network hops; NVMe OSDs are usually CPU-bound, which is the motivation for Crimson.
- Small-block EC was historically slow because of read-modify-write; FastEC with a 16 KiB+ stripe unit is the Tentacle answer.
- Recovery traffic competes with client I/O. Since Quincy the
mclock_scheduler(defaultosd_op_queue) balances client, recovery, and background work using per-OSD IOPS capacity; Tentacle adds sanity thresholds for unrealistically low measured IOPS. - Scale: CERN's "Big Bang III" test ran a stable cluster of more than 10,000 OSDs on a Luminous release candidate.
Indicative throughput and IOPS ranges, with their caveats, are in Reference. They are unsourced estimates; run rados bench and fio on your own hardware.
Governance and Release Model¶
- Origins: Ceph began as Sage Weil's PhD work at UC Santa Cruz. Inktank commercialized it and was acquired by Red Hat in 2014. Red Hat's Ceph storage team moved to IBM in 2023, and IBM is now the largest corporate contributor, alongside companies such as CLYSO, Canonical, and others.
- Ceph Foundation: formed in 2018 as a directed fund under the Linux Foundation. Its Governing Board (all Premier members plus representatives of General and Associate members and of the technical leadership) manages budget, outreach, and events but has no direct control over technical decisions. The Q1 2026 community newsletter reports that the Foundation began a transition within the Linux Foundation in September 2025 and approved a new Foundation Charter that separates funding from technical governance.
- Technical governance: the Ceph Steering Committee (open membership by supermajority vote) sets direction and elects a three-person Ceph Executive Council for one-year terms; council members must come from more than one employer.
- Licence: LGPL-2.1 or LGPL-3.0 for the core, with some components under other licences.
- Releases: one named stable release per year, alphabetical and cephalopod-themed (Reef, Squid, Tentacle, then Umbrella v21). Each stable series gets point releases for roughly two years; the upstream docs list exactly two active series at a time. Upgrades are supported from the two previous stable releases (Tentacle accepts upgrades from Reef or Squid).
Sources¶
- Ceph architecture
- CRUSH maps and the CRUSH paper: Weil et al., "CRUSH: Controlled, Scalable, Decentralized Placement of Replicated Data" (SC 2006)
- BlueStore configuration reference
- Erasure code (EC optimizations)
- Tentacle release notes and Squid release notes
- Crimson (tech preview) and Crimson Tentacle update blog
- NVMe-oF gateway overview
- CVE-2025-30156 and security advisories index
- CephX config reference
- User management
- ceph-volume encryption
- Messenger v2
- Ceph Foundation and Governance
- Ceph Q1 2026 newsletter
- Ceph blog: New in Luminous, improved scalability (CERN Big Bang III)