How-to Guides¶
Scope
Task recipes for running Longhorn v1.12.x on Kubernetes v1.25+: installing, upgrading, provisioning volumes (RWO, RWX, V2), encryption, backup targets, recurring jobs, system backup, DR, tuning and troubleshooting. Setting defaults, ports and URL formats are in Reference. The reasons behind these steps are in Explanation.
Install Longhorn¶
Check and install node prerequisites¶
Every node needs open-iscsi (V1), an NFSv4 client (RWX and NFS backups), cryptsetup and dmsetup. longhornctl checks for them and can install them.
# Download the CLI that matches the Longhorn version (amd64 shown, arm64 also published)
curl -sSfL -o longhornctl \
https://github.com/longhorn/cli/releases/download/v1.12.1/longhornctl-linux-amd64
chmod +x longhornctl
# Check nodes (add --enable-spdk to also check V2 requirements)
./longhornctl check preflight
# Install missing packages and modules on all nodes
./longhornctl --kubeconfig ~/.kube/config \
--image longhornio/longhorn-cli:v1.12.1 install preflight
open-iscsi 2.1.12
Do not install or upgrade hosts to open-iscsi 2.1.12. It breaks V1 volume attachment. Use 2.1.11 or earlier, or 2.1.13 or later.
Install with Helm (recommended)¶
helm repo add longhorn https://charts.longhorn.io
helm repo update
helm install longhorn longhorn/longhorn \
--namespace longhorn-system --create-namespace \
--version 1.12.1 \
--set persistence.defaultClassReplicaCount=3 \
--set defaultSettings.storageOverProvisioningPercentage=200
Other supported methods: kubectl apply -f https://raw.githubusercontent.com/longhorn/longhorn/v1.12.1/deploy/longhorn.yaml, the Rancher Apps catalog, the Helm Controller, Fleet, Flux and Argo CD. Platform-specific steps exist for GKE, K3s, RKE with CoreOS, OKD/OpenShift, Talos and Container-Optimized OS.
Default settings format
Since v1.10, per-engine settings take JSON, for example defaultSettings.defaultReplicaCount='{"v1":"3","v2":"3"}'. A plain value applies to all engines only for settings defined as single-value. Check each setting's format in the settings reference. The persistence.defaultClassReplicaCount value sets numberOfReplicas on the default longhorn StorageClass.
Verify the installation and open the UI¶
kubectl -n longhorn-system get pods
kubectl -n longhorn-system get nodes.longhorn.io -o wide
# Temporary UI access (the UI has no built-in auth, so use an authenticated ingress for real use)
kubectl port-forward -n longhorn-system svc/longhorn-frontend 8080:80
Upgrade Longhorn¶
- Upgrade one minor version at a time (1.11.x to 1.12.x). Skipped minors and downgrades are rejected.
- Before you start: make sure no volume is
Faultedand no BackingImage has failed, detach all V2 volumes, and take a system backup. - For 1.12.x: migrate V2 volumes that use backing images before upgrading (see the v1.12.1 important notes). Legacy V2 linked clones become detach/delete-only.
-
Upgrade the manager:
-
Upgrade engine images, either automatically (Concurrent Automatic Engine Upgrade Per Node Limit > 0) or per volume in the UI. V1 engines upgrade live.
- On a CNI that enforces NetworkPolicy, confirm that the new internal policies (v1.12.1) are not blocking traffic. If needed, retry with
--set networkPolicies.restrictInternalTraffic=false.
# Old instance-manager pods stay until their volumes move. This is expected.
kubectl -n longhorn-system get instancemanagers.longhorn.io
kubectl -n longhorn-system get engineimages.longhorn.io
Provision Volumes¶
StorageClass configuration¶
Do not edit the default longhorn StorageClass, because an upgrade can reset it. Create your own:
kind: StorageClass
apiVersion: storage.k8s.io/v1
metadata:
name: longhorn-fast
provisioner: driver.longhorn.io
allowVolumeExpansion: true
reclaimPolicy: Delete
volumeBindingMode: WaitForFirstConsumer
parameters:
numberOfReplicas: "3"
staleReplicaTimeout: "2880" # minutes (48 h)
dataLocality: "best-effort"
diskSelector: "ssd" # only disks tagged "ssd"
fsType: "ext4"
PersistentVolumeClaim¶
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: mydata
spec:
accessModes:
- ReadWriteOnce # ReadWriteOncePod is also supported since v1.11
storageClassName: longhorn-fast
resources:
requests:
storage: 10Gi
RWX (shared filesystem) volume¶
Every node that mounts the volume needs an NFSv4.1 client. Longhorn starts a share-manager-<volume> pod that serves the volume over NFS.
kind: StorageClass
apiVersion: storage.k8s.io/v1
metadata:
name: longhorn-rwx
provisioner: driver.longhorn.io
allowVolumeExpansion: true
parameters:
numberOfReplicas: "3"
shareManagerNodeSelector: "storage:true" # optional placement for share-manager pods
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: shared-data
spec:
accessModes: ["ReadWriteMany"]
storageClassName: longhorn-rwx
resources:
requests:
storage: 20Gi
kubectl -n longhorn-system get sharemanagers.longhorn.io
kubectl -n longhorn-system get pods -l longhorn.io/component=share-manager
Enable the V2 Data Engine¶
-
Prepare each V2 node: kernel 6.7+ recommended, modules loaded, and 2 GiB of huge pages.
-
Enable the engine (instance-manager pods restart). Make sure the Guaranteed Instance Manager CPU covers the cores in
data-engine-cpu-mask(default0x3, 2 cores). -
Add a raw block-type disk to each node with
kubectl edit nodes.longhorn.io <node>: -
Create a V2 StorageClass:
ARM64 and burstable VMs
On ARM64, use diskDriver: aio disks (NVMe-driver disks can hang with 2+ SPDK cores). On AWS T-series or other burstable VMs, set unlimited CPU credits. A throttled spdk_tgt stalls V2 I/O.
Set Up Volume Encryption¶
-
Create the key Secret. The CSI sidecars must be able to read it.
-
Create a StorageClass with a global encryption key (one key for all volumes):
kind: StorageClass apiVersion: storage.k8s.io/v1 metadata: name: longhorn-crypto-global provisioner: driver.longhorn.io allowVolumeExpansion: true parameters: numberOfReplicas: "3" encrypted: "true" csi.storage.k8s.io/provisioner-secret-name: "longhorn-crypto" csi.storage.k8s.io/provisioner-secret-namespace: "longhorn-system" csi.storage.k8s.io/node-publish-secret-name: "longhorn-crypto" csi.storage.k8s.io/node-publish-secret-namespace: "longhorn-system" csi.storage.k8s.io/node-stage-secret-name: "longhorn-crypto" csi.storage.k8s.io/node-stage-secret-namespace: "longhorn-system" csi.storage.k8s.io/node-expand-secret-name: "longhorn-crypto" csi.storage.k8s.io/node-expand-secret-namespace: "longhorn-system" -
Or use a per-volume encryption key (unique key per PVC). Create a Secret named after each PVC, in the PVC's namespace:
kind: StorageClass apiVersion: storage.k8s.io/v1 metadata: name: longhorn-crypto-per-volume provisioner: driver.longhorn.io allowVolumeExpansion: true parameters: numberOfReplicas: "3" encrypted: "true" csi.storage.k8s.io/provisioner-secret-name: ${pvc.name} csi.storage.k8s.io/provisioner-secret-namespace: ${pvc.namespace} csi.storage.k8s.io/node-publish-secret-name: ${pvc.name} csi.storage.k8s.io/node-publish-secret-namespace: ${pvc.namespace} csi.storage.k8s.io/node-stage-secret-name: ${pvc.name} csi.storage.k8s.io/node-stage-secret-namespace: ${pvc.namespace} csi.storage.k8s.io/node-expand-secret-name: ${pvc.name} csi.storage.k8s.io/node-expand-secret-namespace: ${pvc.namespace}
A PVC stays Pending until its Secret exists. Back up the Secrets separately: restoring an encrypted backup in another cluster needs the same key.
Configure a Backup Target¶
-
Create the credential Secret in
longhorn-system(S3 example): -
Point the
defaultBackupTarget at it. Since v1.8 it is abackuptargets.longhorn.ioCR, not a setting:kubectl -n longhorn-system patch backuptargets.longhorn.io default --type=merge \ -p '{"spec":{"backupTargetURL":"s3://longhorn-backup@us-east-1/","credentialSecret":"aws-secret","pollInterval":"300s"}}' kubectl -n longhorn-system get backuptargets.longhorn.ioWith Helm, set
defaultBackupStore.backupTarget,defaultBackupStore.backupTargetCredentialSecretanddefaultBackupStore.pollIntervalinstead. Other URL schemes:nfs://,cifs://,azblob://(see Reference).
No bucket lifecycle rules
Do not configure S3 lifecycle or retention rules on the backup bucket. Longhorn manages block garbage collection itself, and deleting blocks underneath it corrupts backups.
Schedule Recurring Snapshots and Backups¶
A RecurringJob applies to volumes in its groups. The default group covers every volume that has no explicit recurring-job labels.
apiVersion: longhorn.io/v1beta2
kind: RecurringJob
metadata:
name: daily-backup
namespace: longhorn-system
spec:
cron: "0 2 * * *"
task: backup # snapshot | snapshot-force-create | snapshot-cleanup | snapshot-delete | backup | backup-force-create | filesystem-trim | system-backup
groups:
- default
retain: 7
concurrency: 2
parameters:
full-backup-interval: "7" # optional: every 7th backup is a full backup
# Attach a specific job to one volume instead of using groups
kubectl -n longhorn-system label volumes.longhorn.io <volume> \
recurring-job.longhorn.io/daily-backup=enabled
Create an On-Demand Backup¶
Use the CSI snapshot API with a Longhorn VolumeSnapshotClass of type: bak. This requires the external-snapshotter CRDs and controller.
kind: VolumeSnapshotClass
apiVersion: snapshot.storage.k8s.io/v1
metadata:
name: longhorn-backup-vsc
driver: driver.longhorn.io
deletionPolicy: Delete
parameters:
type: bak # "snap" keeps an in-cluster snapshot only
---
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
name: mydata-backup-1
spec:
volumeSnapshotClassName: longhorn-backup-vsc
source:
persistentVolumeClaimName: mydata
Back Up the Longhorn System¶
apiVersion: longhorn.io/v1beta2
kind: SystemBackup
metadata:
name: pre-upgrade-1-12-1
namespace: longhorn-system
spec:
volumeBackupPolicy: if-not-present # always | if-not-present | disabled
For a schedule, create a RecurringJob with task: system-backup and parameters: {volume-backup-policy: if-not-present}. System backups always go to the default backup target.
Set Up a DR (Standby) Volume¶
- In the DR cluster, point the
defaultBackupTarget at the same backupstore as the primary cluster. - In the UI, go to Backup and Restore > Backups, select the backup volume, and choose Create Disaster Recovery Volume.
- Tune RTO with the backup target
pollInterval. Tune RPO with the source volume's recurring backup schedule. - On failover, Activate the DR volume. Wait until the last backup is restored, because activation fails otherwise. Then create a PV/PVC for it.
Tune for Performance¶
- Use SSD/NVMe, a dedicated disk (not the root disk), and 10 Gbps networking. Put replica traffic on a dedicated storage network (Multus NetworkAttachmentDefinition in the
storage-networksetting). - Set the StorageClass
dataLocality: best-effort. For self-replicating apps (Kafka, Cassandra, CockroachDB), usestrict-localwithnumberOfReplicas: "1"(V1 only). - Use 2 replicas for data-intensive workloads when you accept lower redundancy. Use storage tags (
diskSelector,nodeSelector) for tiering. - Size the Guaranteed Instance Manager CPU. The default is 12% of allocatable CPU per instance-manager pod. Raise it on nodes with many volumes. For V2, it must cover the SPDK cores.
- Keep Replica Node Level Soft Anti-Affinity at
false(the default), so one node failure never takes two replicas of the same volume. - Schedule
snapshot-cleanup/filesystem-trimrecurring jobs to cap snapshot space. - For V2: enable
data-engine-cpu-isolation-enabledso IRQs and RPS stay off SPDK cores. Consider interrupt mode on lightly loaded nodes. - Measure with kbench, the upstream fio-based tool:
kubectl apply -f https://raw.githubusercontent.com/longhorn/kbench/main/deploy/fio.yaml.
Troubleshooting¶
Commands¶
# Volume, engine and replica state
kubectl -n longhorn-system get volumes.longhorn.io
kubectl -n longhorn-system get engines.longhorn.io
kubectl -n longhorn-system get replicas.longhorn.io
# Who is holding a volume attached (attachment tickets, v1.5+)
kubectl -n longhorn-system get volumeattachments.longhorn.io <volume> -o yaml
# Manager and data-plane logs (engine/replica logs live in instance-manager pods)
kubectl -n longhorn-system logs -l app=longhorn-manager --tail=200
kubectl -n longhorn-system logs -l longhorn.io/component=instance-manager,longhorn.io/node=<node> --tail=200
# Node and disk health
kubectl -n longhorn-system get nodes.longhorn.io -o wide
./longhornctl check preflight
# NetworkPolicies created by v1.12.1 (check these when volumes get stuck after an upgrade)
kubectl -n longhorn-system get networkpolicy
Stuck attachments
Older notes suggested patching spec.nodeID on the Volume CR to force a detach. Since v1.5, attachment is driven by tickets in volumeattachments.longhorn.io. Find the stale ticket owner (a deleted pod, VolumeAttachment or backup) and remove it at the source. Do not patch Volume CR fields by hand.
Common Issues¶
| Issue | Diagnosis | Fix |
|---|---|---|
| Volume degraded | UI > Volume details. kubectl get replicas.longhorn.io |
Check node and disk health. Rebuild starts automatically when a schedulable disk exists |
| Rebuild slow | Replica rebuild progress in the UI. longhornctl check preflight for node issues |
Raise Concurrent Replica Rebuild Per Node Limit (default 5). Check the storage network. Fast and parallel rebuild need v1.11+ |
| PVC stuck Pending | kubectl describe pvc. Missing encryption Secret? |
Check the StorageClass, disk tags and CSIStorageCapacity on compute-only nodes (fixed in v1.12.0) |
| Volume stuck attaching after upgrade to 1.12.1 | kubectl -n longhorn-system get networkpolicy. CNI enforcing? |
Adjust or disable internal policies: networkPolicies.restrictInternalTraffic=false |
RWX mount fails with a mount.nfs helper error |
NFS client missing on the node | Install nfs-common / nfs-utils / nfs-client |
| RWX fails with "protocol not supported" | Known-bad kernel (6.5.6, Ubuntu 5.15.0-94, 6.5.0-21, 6.5.0-1014-aws) | Change the kernel version |
| V1 attach fails after an OS update | open-iscsi version |
Avoid 2.1.12 |
| Disk pressure | kubectl -n longhorn-system get nodes.longhorn.io -o yaml (disk status) |
Add disks, clean snapshots, trim filesystems, delete unused volumes |
| V2 volumes slow to attach at scale | Many attached V2 volumes | Known issue #13241 in v1.12.1. Track upstream |
Sources¶
- Quick Installation and requirements
- Install with Helm
- Upgrade
- Command Line Tool (longhornctl)
- ReadWriteMany (RWX) Volume
- Multiple disks and block-type disks
- Volume Encryption
- Setting a Backup Target
- Recurring Snapshots and Backups
- CSI VolumeSnapshot associated with Longhorn backup
- Backup Longhorn System
- Best Practices
- Longhorn CLI / kubectl references