Skip to content

How-to Guides

Scope

Task recipes for running Longhorn v1.12.x on Kubernetes v1.25+: installing, upgrading, provisioning volumes (RWO, RWX, V2), encryption, backup targets, recurring jobs, system backup, DR, tuning and troubleshooting. Setting defaults, ports and URL formats are in Reference. The reasons behind these steps are in Explanation.

Install Longhorn

Check and install node prerequisites

Every node needs open-iscsi (V1), an NFSv4 client (RWX and NFS backups), cryptsetup and dmsetup. longhornctl checks for them and can install them.

# Download the CLI that matches the Longhorn version (amd64 shown, arm64 also published)
curl -sSfL -o longhornctl \
  https://github.com/longhorn/cli/releases/download/v1.12.1/longhornctl-linux-amd64
chmod +x longhornctl

# Check nodes (add --enable-spdk to also check V2 requirements)
./longhornctl check preflight

# Install missing packages and modules on all nodes
./longhornctl --kubeconfig ~/.kube/config \
  --image longhornio/longhorn-cli:v1.12.1 install preflight

open-iscsi 2.1.12

Do not install or upgrade hosts to open-iscsi 2.1.12. It breaks V1 volume attachment. Use 2.1.11 or earlier, or 2.1.13 or later.

helm repo add longhorn https://charts.longhorn.io
helm repo update

helm install longhorn longhorn/longhorn \
  --namespace longhorn-system --create-namespace \
  --version 1.12.1 \
  --set persistence.defaultClassReplicaCount=3 \
  --set defaultSettings.storageOverProvisioningPercentage=200

Other supported methods: kubectl apply -f https://raw.githubusercontent.com/longhorn/longhorn/v1.12.1/deploy/longhorn.yaml, the Rancher Apps catalog, the Helm Controller, Fleet, Flux and Argo CD. Platform-specific steps exist for GKE, K3s, RKE with CoreOS, OKD/OpenShift, Talos and Container-Optimized OS.

Default settings format

Since v1.10, per-engine settings take JSON, for example defaultSettings.defaultReplicaCount='{"v1":"3","v2":"3"}'. A plain value applies to all engines only for settings defined as single-value. Check each setting's format in the settings reference. The persistence.defaultClassReplicaCount value sets numberOfReplicas on the default longhorn StorageClass.

Verify the installation and open the UI

kubectl -n longhorn-system get pods
kubectl -n longhorn-system get nodes.longhorn.io -o wide

# Temporary UI access (the UI has no built-in auth, so use an authenticated ingress for real use)
kubectl port-forward -n longhorn-system svc/longhorn-frontend 8080:80

Upgrade Longhorn

  1. Upgrade one minor version at a time (1.11.x to 1.12.x). Skipped minors and downgrades are rejected.
  2. Before you start: make sure no volume is Faulted and no BackingImage has failed, detach all V2 volumes, and take a system backup.
  3. For 1.12.x: migrate V2 volumes that use backing images before upgrading (see the v1.12.1 important notes). Legacy V2 linked clones become detach/delete-only.
  4. Upgrade the manager:

    helm repo update
    helm upgrade longhorn longhorn/longhorn -n longhorn-system \
      --version 1.12.1 --reuse-values
    
  5. Upgrade engine images, either automatically (Concurrent Automatic Engine Upgrade Per Node Limit > 0) or per volume in the UI. V1 engines upgrade live.

  6. On a CNI that enforces NetworkPolicy, confirm that the new internal policies (v1.12.1) are not blocking traffic. If needed, retry with --set networkPolicies.restrictInternalTraffic=false.
# Old instance-manager pods stay until their volumes move. This is expected.
kubectl -n longhorn-system get instancemanagers.longhorn.io
kubectl -n longhorn-system get engineimages.longhorn.io

Provision Volumes

StorageClass configuration

Do not edit the default longhorn StorageClass, because an upgrade can reset it. Create your own:

kind: StorageClass
apiVersion: storage.k8s.io/v1
metadata:
  name: longhorn-fast
provisioner: driver.longhorn.io
allowVolumeExpansion: true
reclaimPolicy: Delete
volumeBindingMode: WaitForFirstConsumer
parameters:
  numberOfReplicas: "3"
  staleReplicaTimeout: "2880"   # minutes (48 h)
  dataLocality: "best-effort"
  diskSelector: "ssd"           # only disks tagged "ssd"
  fsType: "ext4"

PersistentVolumeClaim

apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: mydata
spec:
  accessModes:
    - ReadWriteOnce            # ReadWriteOncePod is also supported since v1.11
  storageClassName: longhorn-fast
  resources:
    requests:
      storage: 10Gi

RWX (shared filesystem) volume

Every node that mounts the volume needs an NFSv4.1 client. Longhorn starts a share-manager-<volume> pod that serves the volume over NFS.

kind: StorageClass
apiVersion: storage.k8s.io/v1
metadata:
  name: longhorn-rwx
provisioner: driver.longhorn.io
allowVolumeExpansion: true
parameters:
  numberOfReplicas: "3"
  shareManagerNodeSelector: "storage:true"   # optional placement for share-manager pods
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: shared-data
spec:
  accessModes: ["ReadWriteMany"]
  storageClassName: longhorn-rwx
  resources:
    requests:
      storage: 20Gi
kubectl -n longhorn-system get sharemanagers.longhorn.io
kubectl -n longhorn-system get pods -l longhorn.io/component=share-manager

Enable the V2 Data Engine

  1. Prepare each V2 node: kernel 6.7+ recommended, modules loaded, and 2 GiB of huge pages.

    modprobe vfio_pci && modprobe uio_pci_generic && modprobe nvme-tcp
    echo 1024 > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages   # not persistent, set hugepages=1024 on the kernel cmdline too
    systemctl restart kubelet
    ./longhornctl check preflight --enable-spdk
    
  2. Enable the engine (instance-manager pods restart). Make sure the Guaranteed Instance Manager CPU covers the cores in data-engine-cpu-mask (default 0x3, 2 cores).

    kubectl -n longhorn-system patch settings.longhorn.io v2-data-engine \
      --type=merge -p '{"value":"true"}'
    
  3. Add a raw block-type disk to each node with kubectl edit nodes.longhorn.io <node>:

    spec:
      disks:
        nvme-disk:
          allowScheduling: true
          diskType: block
          diskDriver: auto        # nvme via vfio-pci when possible, otherwise aio
          path: /dev/nvme1n1      # prefer a stable /dev/disk/by-id path
          storageReserved: 0
          tags: []
    
  4. Create a V2 StorageClass:

    kind: StorageClass
    apiVersion: storage.k8s.io/v1
    metadata:
      name: longhorn-v2
    provisioner: driver.longhorn.io
    allowVolumeExpansion: true
    parameters:
      dataEngine: "v2"
      numberOfReplicas: "3"
      staleReplicaTimeout: "2880"
      fsType: "ext4"
    

ARM64 and burstable VMs

On ARM64, use diskDriver: aio disks (NVMe-driver disks can hang with 2+ SPDK cores). On AWS T-series or other burstable VMs, set unlimited CPU credits. A throttled spdk_tgt stalls V2 I/O.

Set Up Volume Encryption

  1. Create the key Secret. The CSI sidecars must be able to read it.

    apiVersion: v1
    kind: Secret
    metadata:
      name: longhorn-crypto
      namespace: longhorn-system
    stringData:
      CRYPTO_KEY_VALUE: "<passphrase>"
      CRYPTO_KEY_PROVIDER: "secret"
      CRYPTO_KEY_CIPHER: "aes-xts-plain64"
      CRYPTO_KEY_HASH: "sha256"
      CRYPTO_KEY_SIZE: "256"
      CRYPTO_PBKDF: "argon2i"
    
  2. Create a StorageClass with a global encryption key (one key for all volumes):

    kind: StorageClass
    apiVersion: storage.k8s.io/v1
    metadata:
      name: longhorn-crypto-global
    provisioner: driver.longhorn.io
    allowVolumeExpansion: true
    parameters:
      numberOfReplicas: "3"
      encrypted: "true"
      csi.storage.k8s.io/provisioner-secret-name: "longhorn-crypto"
      csi.storage.k8s.io/provisioner-secret-namespace: "longhorn-system"
      csi.storage.k8s.io/node-publish-secret-name: "longhorn-crypto"
      csi.storage.k8s.io/node-publish-secret-namespace: "longhorn-system"
      csi.storage.k8s.io/node-stage-secret-name: "longhorn-crypto"
      csi.storage.k8s.io/node-stage-secret-namespace: "longhorn-system"
      csi.storage.k8s.io/node-expand-secret-name: "longhorn-crypto"
      csi.storage.k8s.io/node-expand-secret-namespace: "longhorn-system"
    
  3. Or use a per-volume encryption key (unique key per PVC). Create a Secret named after each PVC, in the PVC's namespace:

    kind: StorageClass
    apiVersion: storage.k8s.io/v1
    metadata:
      name: longhorn-crypto-per-volume
    provisioner: driver.longhorn.io
    allowVolumeExpansion: true
    parameters:
      numberOfReplicas: "3"
      encrypted: "true"
      csi.storage.k8s.io/provisioner-secret-name: ${pvc.name}
      csi.storage.k8s.io/provisioner-secret-namespace: ${pvc.namespace}
      csi.storage.k8s.io/node-publish-secret-name: ${pvc.name}
      csi.storage.k8s.io/node-publish-secret-namespace: ${pvc.namespace}
      csi.storage.k8s.io/node-stage-secret-name: ${pvc.name}
      csi.storage.k8s.io/node-stage-secret-namespace: ${pvc.namespace}
      csi.storage.k8s.io/node-expand-secret-name: ${pvc.name}
      csi.storage.k8s.io/node-expand-secret-namespace: ${pvc.namespace}
    

A PVC stays Pending until its Secret exists. Back up the Secrets separately: restoring an encrypted backup in another cluster needs the same key.

Configure a Backup Target

  1. Create the credential Secret in longhorn-system (S3 example):

    kubectl create secret generic aws-secret -n longhorn-system \
      --from-literal=AWS_ACCESS_KEY_ID=<key-id> \
      --from-literal=AWS_SECRET_ACCESS_KEY=<secret>
    # For MinIO or another S3-compatible store, also add:
    #   --from-literal=AWS_ENDPOINTS=https://minio.example.internal:9000
    
  2. Point the default BackupTarget at it. Since v1.8 it is a backuptargets.longhorn.io CR, not a setting:

    kubectl -n longhorn-system patch backuptargets.longhorn.io default --type=merge \
      -p '{"spec":{"backupTargetURL":"s3://longhorn-backup@us-east-1/","credentialSecret":"aws-secret","pollInterval":"300s"}}'
    kubectl -n longhorn-system get backuptargets.longhorn.io
    

    With Helm, set defaultBackupStore.backupTarget, defaultBackupStore.backupTargetCredentialSecret and defaultBackupStore.pollInterval instead. Other URL schemes: nfs://, cifs://, azblob:// (see Reference).

No bucket lifecycle rules

Do not configure S3 lifecycle or retention rules on the backup bucket. Longhorn manages block garbage collection itself, and deleting blocks underneath it corrupts backups.

Schedule Recurring Snapshots and Backups

A RecurringJob applies to volumes in its groups. The default group covers every volume that has no explicit recurring-job labels.

apiVersion: longhorn.io/v1beta2
kind: RecurringJob
metadata:
  name: daily-backup
  namespace: longhorn-system
spec:
  cron: "0 2 * * *"
  task: backup            # snapshot | snapshot-force-create | snapshot-cleanup | snapshot-delete | backup | backup-force-create | filesystem-trim | system-backup
  groups:
    - default
  retain: 7
  concurrency: 2
  parameters:
    full-backup-interval: "7"   # optional: every 7th backup is a full backup
# Attach a specific job to one volume instead of using groups
kubectl -n longhorn-system label volumes.longhorn.io <volume> \
  recurring-job.longhorn.io/daily-backup=enabled

Create an On-Demand Backup

Use the CSI snapshot API with a Longhorn VolumeSnapshotClass of type: bak. This requires the external-snapshotter CRDs and controller.

kind: VolumeSnapshotClass
apiVersion: snapshot.storage.k8s.io/v1
metadata:
  name: longhorn-backup-vsc
driver: driver.longhorn.io
deletionPolicy: Delete
parameters:
  type: bak            # "snap" keeps an in-cluster snapshot only
---
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
  name: mydata-backup-1
spec:
  volumeSnapshotClassName: longhorn-backup-vsc
  source:
    persistentVolumeClaimName: mydata
kubectl get volumesnapshot mydata-backup-1
kubectl -n longhorn-system get backups.longhorn.io

Back Up the Longhorn System

apiVersion: longhorn.io/v1beta2
kind: SystemBackup
metadata:
  name: pre-upgrade-1-12-1
  namespace: longhorn-system
spec:
  volumeBackupPolicy: if-not-present   # always | if-not-present | disabled
kubectl -n longhorn-system get systembackups.longhorn.io   # wait for STATE Ready

For a schedule, create a RecurringJob with task: system-backup and parameters: {volume-backup-policy: if-not-present}. System backups always go to the default backup target.

Set Up a DR (Standby) Volume

  1. In the DR cluster, point the default BackupTarget at the same backupstore as the primary cluster.
  2. In the UI, go to Backup and Restore > Backups, select the backup volume, and choose Create Disaster Recovery Volume.
  3. Tune RTO with the backup target pollInterval. Tune RPO with the source volume's recurring backup schedule.
  4. On failover, Activate the DR volume. Wait until the last backup is restored, because activation fails otherwise. Then create a PV/PVC for it.

Tune for Performance

  1. Use SSD/NVMe, a dedicated disk (not the root disk), and 10 Gbps networking. Put replica traffic on a dedicated storage network (Multus NetworkAttachmentDefinition in the storage-network setting).
  2. Set the StorageClass dataLocality: best-effort. For self-replicating apps (Kafka, Cassandra, CockroachDB), use strict-local with numberOfReplicas: "1" (V1 only).
  3. Use 2 replicas for data-intensive workloads when you accept lower redundancy. Use storage tags (diskSelector, nodeSelector) for tiering.
  4. Size the Guaranteed Instance Manager CPU. The default is 12% of allocatable CPU per instance-manager pod. Raise it on nodes with many volumes. For V2, it must cover the SPDK cores.
  5. Keep Replica Node Level Soft Anti-Affinity at false (the default), so one node failure never takes two replicas of the same volume.
  6. Schedule snapshot-cleanup / filesystem-trim recurring jobs to cap snapshot space.
  7. For V2: enable data-engine-cpu-isolation-enabled so IRQs and RPS stay off SPDK cores. Consider interrupt mode on lightly loaded nodes.
  8. Measure with kbench, the upstream fio-based tool: kubectl apply -f https://raw.githubusercontent.com/longhorn/kbench/main/deploy/fio.yaml.

Troubleshooting

Commands

# Volume, engine and replica state
kubectl -n longhorn-system get volumes.longhorn.io
kubectl -n longhorn-system get engines.longhorn.io
kubectl -n longhorn-system get replicas.longhorn.io

# Who is holding a volume attached (attachment tickets, v1.5+)
kubectl -n longhorn-system get volumeattachments.longhorn.io <volume> -o yaml

# Manager and data-plane logs (engine/replica logs live in instance-manager pods)
kubectl -n longhorn-system logs -l app=longhorn-manager --tail=200
kubectl -n longhorn-system logs -l longhorn.io/component=instance-manager,longhorn.io/node=<node> --tail=200

# Node and disk health
kubectl -n longhorn-system get nodes.longhorn.io -o wide
./longhornctl check preflight

# NetworkPolicies created by v1.12.1 (check these when volumes get stuck after an upgrade)
kubectl -n longhorn-system get networkpolicy

Stuck attachments

Older notes suggested patching spec.nodeID on the Volume CR to force a detach. Since v1.5, attachment is driven by tickets in volumeattachments.longhorn.io. Find the stale ticket owner (a deleted pod, VolumeAttachment or backup) and remove it at the source. Do not patch Volume CR fields by hand.

Common Issues

Issue Diagnosis Fix
Volume degraded UI > Volume details. kubectl get replicas.longhorn.io Check node and disk health. Rebuild starts automatically when a schedulable disk exists
Rebuild slow Replica rebuild progress in the UI. longhornctl check preflight for node issues Raise Concurrent Replica Rebuild Per Node Limit (default 5). Check the storage network. Fast and parallel rebuild need v1.11+
PVC stuck Pending kubectl describe pvc. Missing encryption Secret? Check the StorageClass, disk tags and CSIStorageCapacity on compute-only nodes (fixed in v1.12.0)
Volume stuck attaching after upgrade to 1.12.1 kubectl -n longhorn-system get networkpolicy. CNI enforcing? Adjust or disable internal policies: networkPolicies.restrictInternalTraffic=false
RWX mount fails with a mount.nfs helper error NFS client missing on the node Install nfs-common / nfs-utils / nfs-client
RWX fails with "protocol not supported" Known-bad kernel (6.5.6, Ubuntu 5.15.0-94, 6.5.0-21, 6.5.0-1014-aws) Change the kernel version
V1 attach fails after an OS update open-iscsi version Avoid 2.1.12
Disk pressure kubectl -n longhorn-system get nodes.longhorn.io -o yaml (disk status) Add disks, clean snapshots, trim filesystems, delete unused volumes
V2 volumes slow to attach at scale Many attached V2 volumes Known issue #13241 in v1.12.1. Track upstream

Sources