How-to Guides¶
Scope
Task recipes for Ceph Tentacle (20.2) and Squid (19.2): deploying with cephadm or Rook, upgrading, migrating CephX keys after the August 2026 CVEs, managing pools, CRUSH, RBD, RGW, and CephFS, securing transport, and troubleshooting. Background is in Explanation; defaults, ports, and version matrices are in Reference.
IP addresses below use the documentation range 192.0.2.0/24; replace them with your own.
Plan a Cluster¶
| Component | Role | Production minimum |
|---|---|---|
| MON | Cluster maps, Paxos quorum, CephX | 3 (odd number; 5 for large or multi-rack) |
| MGR | Dashboard, metrics, orchestrator | 2 (active + standby) |
| OSD | Data storage (one per device) | 3+ hosts, each with OSDs |
| MDS | CephFS metadata (only with CephFS) | 2+ (active + standby) |
| RGW | S3/Swift gateway (only with object) | 2+ behind a load balancer or ingress |
| NVMe-oF gateway | NVMe/TCP block export (optional) | 2+ in one gateway group for HA |
Use replicated pools with size=3, min_size=2 for data you care about. size=2 is technically possible, but the upstream docs warn that two-copy pools eventually lose data through overlapping failures.
Deploy a Cluster with cephadm¶
cephadm is the upstream-recommended installer outside Kubernetes. It needs Python 3, systemd, and Podman or Docker on every host.
# Install cephadm from your distro packages (or the Ceph repo), then bootstrap the first host
cephadm bootstrap --mon-ip 192.0.2.11 --initial-dashboard-user admin
# Distribute the cluster SSH key and add hosts
ssh-copy-id -f -i /etc/ceph/ceph.pub root@node2
ceph orch host add node2 192.0.2.12
ceph orch host add node3 192.0.2.13
# Create OSDs on every unused, unpartitioned device
ceph orch apply osd --all-available-devices
# Optional services
ceph orch apply rgw myrgw --placement="count:2"
ceph fs volume create myfs # also deploys MDS daemons
ceph orch ls # verify services
Encrypt OSDs from the start
For dm-crypt OSDs, apply an OSD service spec with encrypted: true instead of --all-available-devices. Encryption cannot be added to an existing OSD in place; OSDs must be redeployed.
Deploy Ceph on Kubernetes with Rook¶
Rook v1.20.7 (current stable, 2026-09) supports Kubernetes v1.31 to v1.37 and Ceph Squid and Tentacle.
# Operator via Helm
helm repo add rook-release https://charts.rook.io/release
helm install --create-namespace --namespace rook-ceph rook-ceph rook-release/rook-ceph
# Cluster via the rook-ceph-cluster chart (image defaults to quay.io/ceph/ceph:v20.2.4)
helm install --namespace rook-ceph rook-ceph-cluster \
--set operatorNamespace=rook-ceph rook-release/rook-ceph-cluster
# Or with the example manifests from a tagged release
git clone --single-branch --branch v1.20.7 https://github.com/rook/rook.git
cd rook/deploy/examples
kubectl create -f crds.yaml -f common.yaml -f csi-operator.yaml
kubectl create -f operator.yaml -f cluster.yaml
# Check health from the toolbox
kubectl -n rook-ceph exec deploy/rook-ceph-tools -- ceph status
Upgrade a cephadm Cluster¶
Tentacle accepts upgrades from Reef or Squid. Reef is EOL and Squid reaches its target EOL on 2026-10-31, so plan the move to Tentacle now.
# Pre-flight: cluster must be HEALTH_OK with no down/incomplete/backfilling PGs
ceph health detail
ceph osd pool set noautoscale # optional: pause the PG autoscaler
# Start, watch, pause, or stop the rolling upgrade
ceph orch upgrade start --image quay.io/ceph/ceph:v20.2.4
ceph orch upgrade status
ceph -W cephadm
ceph orch upgrade pause # or: resume / stop
# Afterwards
ceph versions
ceph osd pool unset noautoscale
No downgrade
Stopping an upgrade only halts it. There is no path back to Reef or Squid once monitors run Tentacle. On package-based (non-cephadm) clusters, upgrade in order: monitors, managers, OSDs, MDS, then RGW and other gateways, and set noout for the duration.
Rotate CephX Keys to aes256k¶
After upgrading to 20.2.4 or 19.2.6 (fixing CVE-2025-30156), migrate every key to the new aes256k type. cephadm automates daemon keys but not client keys, and Rook automates some client keys. This is the manual procedure from the upstream CephX reference, condensed.
# 1. Confirm the monitors accept the new cipher
ceph --format=json mon dump | jq -r '.auth_allowed_ciphers | map(.name) | join(",")' # expect aes,aes256k
ceph mon set auth_allowed_ciphers aes,aes256k # only if missing
# 2. Make new keys use aes256k by default
ceph mon set auth_preferred_cipher aes256k
# 3. Rotate service daemon keys, starting with mon. (keep the output safe)
ceph auth rotate --key-type=aes256k mon. | tee mon.keyring
systemctl restart ceph-mon@$ID # each monitor
# For each mgr/osd/mds: stop it, rotate, import, restart
systemctl stop ceph-osd@$ID && ceph osd down $ID
ceph auth rotate --key-type=aes256k osd.$ID | tee keyring
ceph-authtool --import-keyring keyring /var/lib/ceph/osd/ceph-$ID/keyring
systemctl restart ceph-osd@$ID
# 4. Switch rotating service keys and block new insecure keys
ceph mon set auth_service_cipher aes256k
ceph config set mon mon_auth_allow_insecure_key false
# 5. Rotate client.admin (create a backup admin key first), then other clients
ceph auth get-or-create client.admin-backup mon "allow *" | tee ./client.admin-backup.keyring
ceph auth rotate --key-type=aes256k client.admin | tee ./client.admin.keyring
ceph-authtool --import-keyring ./client.admin.keyring /etc/ceph/ceph.client.admin.keyring
ceph auth rm client.admin-backup
ceph auth rotate --key-type=aes256k client.$NAME | tee ./client.$NAME.keyring
# 6. Confirm the AUTH_INSECURE_* health checks are gone
ceph health detail | grep AUTH_INSECURE
Before rotating client keys
Kernel clients (krbd, CephFS kernel mounts, Ceph-CSI on Kubernetes nodes) need Linux 7.0+ or a vendor backport to use aes256k keys. Rotate a client key only after every machine using it runs a capable kernel or userspace client. OSDs created by ceph-volume may also need their BlueStore osd_key label updated with ceph-bluestore-tool set-label-key; see the upstream procedure.
CVE-2026-50152: rotate stored secrets too
Before 20.2.4/19.2.6, any key with mon allow r could read the monitor config-key store, which holds OSD LUKS passphrases and the cephadm SSH private key. Upgrading stops further reads but does not revoke what may already have leaked. Rotate the cephadm SSH key and review other secrets in ceph config-key ls; upstream says formal guidance for rotating all config-key secrets is forthcoming.
Manage Pools¶
# Replicated pool (PG count is managed by the autoscaler)
ceph osd pool create mypool
ceph osd pool set mypool size 3
ceph osd pool set mypool min_size 2
ceph osd pool application enable mypool rbd
# Erasure-coded pool with an explicit profile
ceph osd erasure-code-profile set ec42 k=4 m=2 crush-failure-domain=host
ceph osd pool create ecpool erasure ec42
ceph osd pool set ecpool allow_ec_overwrites true # needed for RBD/CephFS data
# Inline compression
ceph osd pool set mypool compression_algorithm zstd
ceph osd pool set mypool compression_mode aggressive
# PG autoscaler view
ceph osd pool autoscale-status
Enable FastEC on an Erasure-Coded Pool (Tentacle)¶
# All MONs and OSDs must run Tentacle. One-way: cannot be turned off again.
ceph osd pool set ecpool allow_ec_optimizations true
# For new pools, pick a 16K+ stripe unit at creation and make optimizations the default
ceph osd erasure-code-profile set ec42fast k=4 m=2 stripe_unit=16K crush-failure-domain=host
ceph config set global osd_pool_default_flag_ec_optimizations true
ceph osd pool create ecfast erasure ec42fast
Track Pool Availability (Tentacle, tech preview)¶
ceph config set mon enable_availability_tracking true
ceph osd pool availability-status
ceph osd pool clear-availability-status mypool
Manage the CRUSH Map¶
# View the hierarchy
ceph osd tree
ceph osd crush tree --show-shadow
# Replicated rule with rack as failure domain, then attach it to a pool
ceph osd crush rule create-replicated replicated_rack default rack
ceph osd pool set mypool crush_rule replicated_rack
# Place an OSD explicitly
ceph osd crush set osd.5 1.0 root=default datacenter=dc1 rack=rack2 host=node5
# Rule restricted to a device class (e.g., NVMe only)
ceph osd crush rule create-replicated fast default host nvme
Use RBD Block Storage¶
rbd pool init mypool
rbd create mypool/myimage --size 100G
rbd create --size 100G --data-pool ecpool mypool/ecimage # data in EC, metadata replicated
# Map with the kernel client (msgr2 is the default since Tentacle)
rbd device map mypool/myimage
mkfs.xfs /dev/rbd0 && mount /dev/rbd0 /mnt/rbd
# Snapshots
rbd snap create mypool/myimage@snap1
rbd snap rollback mypool/myimage@snap1
Use RGW Object Storage¶
# Create an S3 user and read its keys
radosgw-admin user create --uid=myuser --display-name="My User"
radosgw-admin user info --uid=myuser
# Test with the AWS CLI (cephadm RGW listens on port 80 by default)
aws --endpoint-url=http://rgw.example.com s3 mb s3://mybucket
aws --endpoint-url=http://rgw.example.com s3 cp file.txt s3://mybucket/
Use CephFS¶
ceph fs volume create myfs
ceph fs authorize myfs client.fsuser / rw | tee /etc/ceph/ceph.client.fsuser.keyring
# Kernel mount, current device-string syntax: <user>@<fsid>.<fs_name>=<path>
mount -t ceph fsuser@.myfs=/ /mnt/cephfs -o secretfile=/etc/ceph/fsuser.secret
# FUSE mount
ceph-fuse --id fsuser /mnt/cephfs
The older mount -t ceph mon1:/ /mnt/cephfs -o name=...,secret=... form still works on kernels that support it, but the device-string syntax above is what the current docs use.
Manage Users and Capabilities¶
ceph auth ls
ceph auth get-or-create client.app1 mon 'allow r' osd 'allow rw pool=myapp'
ceph auth get-or-create client.glance mon 'profile rbd' osd 'profile rbd pool=images' mgr 'profile rbd pool=images'
ceph auth caps client.app1 mon 'allow r' osd 'allow rwx pool=myapp'
ceph auth get client.app1
ceph auth rm client.app1
Secure Transport¶
# Require msgr2 secure mode (encrypted) for all connections
ceph config set global ms_cluster_mode secure
ceph config set global ms_service_mode secure
ceph config set global ms_client_mode secure
For a manually configured RGW, TLS terminates in the Beast frontend:
rgw_frontends = beast ssl_port=443 ssl_certificate=/etc/ceph/rgw.crt ssl_private_key=/etc/ceph/rgw.key
On Tentacle, 20.2.3 adds ssl_ciphersuites for TLS 1.3 cipher control, and cephadm's certmgr can issue and rotate RGW certificates.
Monitor Health¶
ceph status
ceph health detail
ceph df
ceph osd df tree
ceph osd perf
ceph pg stat
ceph pg dump_stuck unclean
ceph mgr module enable prometheus # metrics on :9283
Troubleshoot Common Issues¶
| Issue | Diagnosis | Fix |
|---|---|---|
HEALTH_WARN PGs degraded |
ceph health detail |
Wait for recovery or restore the missing OSD |
| OSD down | ceph osd tree, journalctl -u ceph-<fsid>@osd.<id> |
Check disk and network, restart the daemon |
| Slow ops | ceph daemon osd.X dump_ops_in_flight, ceph osd perf |
Check disk latency and network; on EC clusters with slow ops, upstream suggests trying the wpq scheduler |
| Near-full OSDs | ceph osd df |
Add capacity, enable the balancer, delete data |
| Clock skew | ceph time-sync-status |
Fix NTP/chrony on all monitor hosts |
AUTH_INSECURE_* warnings after upgrade |
ceph health detail |
Finish the aes256k key rotation |
| Inconsistent PG | ceph health detail \| grep inconsistent |
rados list-inconsistent-obj <pgid>, then ceph pg repair <pgid> |
Commands & Recipes¶
# Maintenance: stop rebalancing while a host is down
ceph osd set noout
ceph osd unset noout
# Put a host into maintenance with cephadm
ceph orch host maintenance enter node2
ceph orch host maintenance exit node2
# Recovery and CRUSH inspection
ceph pg dump pgs_brief | grep -i recover
ceph osd crush dump | jq '.buckets'
# Replace a failed OSD with cephadm, keeping its ID
ceph orch osd rm 5 --replace
ceph orch osd rm status