Skip to content

Explanation

Longhorn is a cloud-native distributed block storage system for Kubernetes. It implements storage using containers and microservices. Each volume gets a dedicated storage controller (the Longhorn Engine), and the system synchronously replicates data across multiple replicas on separate nodes. This page explains how the pieces fit together and why they are designed this way. It covers both data engines as of v1.12.1: V1 (iSCSI, Linux processes) and V2 (SPDK, GA since v1.12.0).

See also: Longhorn hub, How-to Guides, Reference.

Architecture

High-Level Component Diagram

The diagram shows the real Longhorn components in the longhorn-system namespace and how a Pod's volume I/O reaches replicas on three nodes.

graph TB
    subgraph K8S["Kubernetes cluster"]
        API["kube-apiserver<br/>(Longhorn CRs)"]
        subgraph CP["Control plane (longhorn-system)"]
            CSI["longhorn-csi-plugin<br/>+ CSI sidecars"]
            LHM["longhorn-manager<br/>DaemonSet"]
            UI["longhorn-ui"]
            WH["admission and conversion webhooks"]
        end
        subgraph NA["Node A (workload node)"]
            POD["Application Pod"]
            IMA["instance-manager<br/>engine: iSCSI (V1) or NVMe-TCP (V2)"]
            RA["Replica A"]
        end
        subgraph NB["Node B"]
            IMB["instance-manager"]
            RB["Replica B"]
        end
        subgraph NC["Node C"]
            IMC["instance-manager"]
            RC["Replica C"]
        end
        SM["share-manager-VOLUME<br/>NFS-Ganesha (RWX only)"]
    end
    BS[("Backupstore<br/>S3 / NFS / CIFS / Azure Blob")]

    POD -->|"PVC mount"| CSI
    CSI -->|"REST :9500"| LHM
    UI -->|"REST :9500"| LHM
    LHM -->|"watch/update CRs"| API
    WH --- API
    LHM -->|"gRPC :8500-8504"| IMA
    IMA -->|"block device"| POD
    IMA -->|"sync writes"| RA
    IMA -->|"sync writes :10000-30000"| IMB
    IMA -->|"sync writes :10000-30000"| IMC
    IMB --- RB
    IMC --- RC
    LHM --> SM
    IMA -->|"backup blocks"| BS

Two-Layer Design

Longhorn separates concerns into a control plane and a data plane:

Layer Component Role
Control Plane Longhorn Manager Orchestrates volumes, handles CSI and UI API calls, reconciles CRs
Control Plane CSI plugin and sidecars Translates Kubernetes storage calls into Longhorn API calls
Data Plane Instance Manager Per-node pod (one per data engine) that hosts engine and replica instances
Data Plane Longhorn Engine Per-volume storage controller that replicates synchronously
Data Plane Share Manager Per-RWX-volume NFSv4.1 server
Data Plane Backing Image Manager Distributes VM/base images (V1 only since v1.12.0)

Longhorn Manager

The Longhorn Manager runs as a Kubernetes DaemonSet with one pod per node. It follows the Kubernetes controller/operator pattern:

  1. It watches Longhorn Custom Resources (volumes, engines, replicas, nodes, settings, volumeattachments, and others) through the Kubernetes API server.
  2. When a new Volume CR is created (for example by a CSI CreateVolume call), it schedules replicas on separate nodes. On attach, it asks the Instance Manager on the workload node to start the engine.
  3. It serves the Longhorn REST API (port 9500) used by the UI, the CSI plugin and recurring-job pods.
  4. It manages the lifecycle of Instance Manager, Backing Image Manager and Share Manager pods.

Since v1.12.0, the manager's informer caching was trimmed to cut memory use in clusters with many non-Longhorn pods.

Instance Manager

The Instance Manager is a system-managed pod: one per node per data engine version. It hosts the engine and replica instances for every volume that lands on that node. It replaced the old model of one pod per engine or replica, which cut pod overhead. The manager reaches hosted instances through a proxy service inside the pod, so attach, snapshot, backup and rebuild calls all flow through it.

Design consequences:

  • It is the real failure domain. If an Instance Manager pod restarts or is evicted, every engine and replica it hosts goes with it. Longhorn therefore reserves CPU for it (Guaranteed Instance Manager CPU, default 12%), protects it with a PodDisruptionBudget (Node Drain Policy), and gives it the longhorn-critical PriorityClass.
  • Upgrades run side by side. A new Instance Manager pod starts next to the old one. V1 volumes move only when detached and reattached, or through a live engine upgrade. The old pod stays until it is empty, so nodes need spare capacity for both reservations during an upgrade.
  • Debugging target. There is no per-volume engine pod. Engine and replica logs live in instance-manager-<hash> pods, labelled longhorn.io/node=<node>.

Longhorn Engine: V1 and V2 Data Engines

The Engine is the per-volume storage controller. It always runs on the same node as the Pod that consumes the volume. This keeps the frontend local: the only network hops are engine-to-replica writes.

The flowchart below contrasts the two data paths from the application's block device down to disk.

flowchart LR
    subgraph V1["V1 Data Engine (default)"]
        A1["/dev/longhorn/VOLUME"] --> B1["iSCSI initiator<br/>(host open-iscsi)"]
        B1 --> C1["tgt iSCSI target +<br/>engine process"]
        C1 --> D1["replica processes"]
        D1 --> E1["sparse files on<br/>ext4/XFS disk"]
    end
    subgraph V2["V2 Data Engine (SPDK, GA in v1.12.0)"]
        A2["/dev/longhorn/VOLUME"] --> B2["NVMe-TCP initiator<br/>(or UBLK, experimental)"]
        B2 --> C2["spdk_tgt: RAID bdev<br/>(engine)"]
        C2 --> D2["lvol bdevs (replicas)<br/>in an LVS"]
        D2 --> E2["raw block disk<br/>(NVMe via vfio-pci or AIO)"]
    end
Aspect V1 V2
Status GA (default) GA since v1.12.0 (Technical Preview in v1.11, experimental before)
Frontend iSCSI (needs open-iscsi) NVMe-TCP (default) or UBLK (experimental)
Engine Linux process in the Instance Manager SPDK RAID bdev inside spdk_tgt
Replica Linux process over sparse files SPDK logical volume (lvol) bdev on a block-type disk
CPU model Interrupt-driven kernel I/O Busy-polling reactors (2 cores by default since v1.12.0). Optional interrupt mode (v1.10+)
Memory Page cache and process memory 2 GiB huge pages by default (can be disabled since v1.10)
Live engine upgrade Yes No, detach before upgrade. Planned for 1.12 to 1.13

Why SPDK? The V1 path crosses the kernel several times per I/O (iSCSI initiator, TCP, tgt, the engine process, replica file I/O). V2 moves the target, RAID and logical-volume layers into user space with polled NVMe drivers. That removes context switches and copies, at the cost of dedicated CPU cores and huge pages. Community and upstream benchmarks report much lower latency and higher IOPS than V1 on NVMe. See Reference: Benchmarks and Scalability for caveats.

V1 vs V2 Feature Parity

Parity is close but not complete in v1.12.1. V2 lacks strict-local data locality, offline fast rebuild, the revision counter and engine live upgrade. V2 backing images were removed in v1.12.0 in favour of CDI. V2 has features V1 lacks: QoS and fast linked clones. See the feature matrix.

V2 known issues in v1.12.1

Attach latency grows as more V2 volumes are attached (NVMe-TCP connection handling, issue #13241). ARM64 NVMe-backed disks can hang I/O with 2 or more SPDK cores, so use AIO disks (issue #13243). UBLK can panic kernel 6.17.

Volume Replication

Replica Placement and Synchronous Replication

The Engine fans every write out to all healthy replicas in parallel. It acknowledges the application only after every replica has acknowledged.

sequenceDiagram
    participant App as Application Pod
    participant Eng as Engine (Node A)
    participant R1 as Replica 1 (Node A)
    participant R2 as Replica 2 (Node B)
    participant R3 as Replica 3 (Node C)

    App->>Eng: WRITE block
    par Synchronous replication
        Eng->>R1: write (local)
        Eng->>R2: write (network)
        Eng->>R3: write (network)
    end
    R1-->>Eng: ACK
    R2-->>Eng: ACK
    R3-->>Eng: ACK
    Eng-->>App: write complete
    Note over Eng: A replica that fails a write is marked ERR and dropped from the set

Key properties:

  • Synchronous replication: The Engine waits for all healthy replicas to acknowledge before completing the write. There is no quorum write: a replica that errors is marked failed and removed from the set, so writes continue on the survivors.
  • Failure tolerance: With N replicas, the volume tolerates N-1 replica failures and stays operational.
  • Default replica count: 3 (setting {"v1":"3","v2":"3"}). Override it per StorageClass with the numberOfReplicas parameter.
  • Reads: They can be served by any healthy replica. That is why aggregate read throughput scales with replica count in the official benchmark, while writes pay the full replication cost.
  • Placement: Replica Node Level Soft Anti-Affinity is false by default, so two replicas of one volume never share a node. Zone and disk anti-affinity, storage tags and the balance-aware disk selection (v1.11) refine placement.

Replica Rebuilding

When a replica fails, the Manager builds a replacement without taking the volume offline:

sequenceDiagram
    participant LHM as longhorn-manager
    participant Eng as Engine
    participant Good as Healthy replica(s)
    participant New as New blank replica

    Eng->>LHM: replica X marked ERR
    LHM->>New: create blank replica on another node
    LHM->>Eng: add replica
    Eng->>Eng: pause I/O
    Eng->>New: add in write-only (WO) mode
    Eng->>Good: take system snapshot
    Eng->>Eng: resume I/O (new writes also go to New)
    Good-->>New: sync historical snapshot files in background
    Eng->>LHM: sync done
    LHM->>Eng: set New to RW, remove replica X

Rebuild behaviour has improved release by release. Delta and fast rebuild reuse existing data and snapshot checksums. v1.11 added scale replica rebuilding for V1, which streams from several healthy replicas at once. It also made Offline Replica Rebuilding a single global setting. v1.12.0 fixed an instance-manager panic during rebuild storms (issue #13087). Concurrent rebuilds per node are capped at 5 by default.

Replica Storage: Sparse Files and Read Index

Each V1 replica is stored as a chain of Linux sparse files (differencing disks) on a filesystem-type disk. This gives thin provisioning out of the box: a 1 TB volume holding 10 GB of data consumes only about 10 GB on disk. V2 replicas are SPDK lvols in a Logical Volume Store on a raw block disk, which are thin-provisioned too.

For V1, Longhorn keeps an in-memory read index per replica. It is a byte array with one byte per 4 KB block, recording which snapshot layer or live-data layer holds the latest data for each block. It removes the need to scan the whole snapshot chain on reads. It is also why one volume can have at most 254 snapshots.

  • A 1 TB volume uses about 256 MB of read index memory per replica.
  • Write operations reset the read index entry to point to live data.
  • Read operations look up the index to find the correct source layer.

Snapshots and Consistency

A snapshot freezes the current differencing disk and starts a new one on top of it (the volume head). Snapshots come from users, from recurring jobs, and from the system: every rebuild starts with a system snapshot.

  • Deletion on V1 marks a snapshot as removed. Its data is coalesced later by a purge, and the snapshot directly under the volume head cannot be merged right away. V2 merges the parent into the head live, so deletes take effect at once.
  • Crash consistency: Longhorn runs sync before a snapshot. It is crash-consistent, not application-consistent. The optional Freeze Filesystem For Snapshot setting (V1, kernel 5.17+ recommended) freezes the filesystem during the snapshot. Databases still need their own quiescing hooks for application consistency.
  • Integrity: Snapshot checksums are computed on a schedule (or on demand with longhornctl since v1.12.0). They also speed up fast rebuilds.

CSI Driver Integration

The Longhorn CSI driver (driver.longhorn.io) turns Kubernetes storage calls into Longhorn API calls.

sequenceDiagram
    participant PVC as PersistentVolumeClaim
    participant K8s as kube-apiserver
    participant Prov as csi-provisioner sidecar
    participant CSI as longhorn-csi-plugin
    participant LHM as longhorn-manager
    participant IM as instance-manager

    PVC->>K8s: create PVC (storageClassName longhorn)
    K8s->>Prov: watch event
    Prov->>CSI: CreateVolume
    CSI->>LHM: POST /v1/volumes
    LHM->>K8s: create Volume CR and schedule replicas
    Prov->>K8s: create PV, PVC Bound
    Note over K8s,IM: Pod scheduled, attach via csi-attacher
    CSI->>LHM: attach volume to node
    LHM->>IM: start engine and replicas
    IM-->>CSI: /dev/longhorn/VOLUME ready
    CSI->>CSI: NodeStage (format, mount) and NodePublish (bind mount)

The Longhorn CSI driver handles:

  • CreateVolume / DeleteVolume: Provisioning and teardown.
  • ControllerPublishVolume / ControllerUnpublishVolume: Attach and detach to nodes. Since v1.5, attachment requests are tracked as tickets in volumeattachments.longhorn.io.
  • NodeStageVolume / NodePublishVolume: Format and mount the block device into the Pod.
  • CreateSnapshot / DeleteSnapshot: The VolumeSnapshotClass parameter type selects an in-cluster Longhorn snapshot (snap) or a backup (bak). A third value, bi, creates a Longhorn backing image from the volume (csiSnapshotTypeLonghornBackingImage in longhorn-manager csi/controller_server.go, checked 2026-09-27).
  • ControllerExpandVolume / NodeExpandVolume: Online volume expansion.
  • Storage capacity tracking: CSIStorageCapacity objects (v1.10+) feed WaitForFirstConsumer scheduling. v1.12.0 fixed zero-capacity reports on compute-only nodes.
  • Topology: allowedTopologies (v1.11) plus csi-allowed-topology-keys and strictTopology (v1.12.0) control PV node affinity.

For encrypted volumes, the CSI driver passes the encryption Secret to cryptsetup / dm_crypt on the host (see Volume Encryption).

Share Manager (RWX Volumes)

Longhorn provides ReadWriteMany (RWX) access by putting an NFS server in front of an ordinary Longhorn block volume.

  • Each generic RWX volume gets its own share-manager-<volume> pod in longhorn-system. It runs NFS-Ganesha exporting NFSv4.1, with a Service in front. The volume's engine attaches on the share-manager's node, and the CSI plugin NFS-mounts the export into every consuming Pod.
  • Lock recovery: client identities are stored in the longhorn-recovery-backend service (port 9503). If a share-manager pod dies, Longhorn recreates it, and clients reclaim locks within a 90 s NFS grace period, during which I/O blocks. The experimental RWX Volume Fast Failover setting adds a direct heartbeat and cuts the grace period to 30 s. Node hostnames must be unique, and CoreDNS must be healthy, because the recovery backend is reached through DNS.
  • Migratable RWX volumes are a separate type for KubeVirt/Harvester VM live migration. They require volumeMode: Block and migratable: "true", and do not use NFS.
  • RWX volumes through the share manager do not support volumeMode: Block.
  • Placement is controlled by the StorageClass parameters shareManagerNodeSelector and shareManagerTolerations plus allowedTopologies. v1.11 added an optional extra network interface for the share manager.

Backups and Disaster Recovery

Backup Architecture

A backup is a flattened copy of one snapshot, stored outside the cluster in a backupstore.

flowchart LR
    subgraph Primary["Primary storage (replica)"]
        S1["snap1"] --> S2["snap2"] --> S3["snap3"] --> HEAD["volume-head"]
    end
    subgraph Store["Backupstore: backupstore/volumes/.../VOLUME"]
        CFG["volume.cfg"]
        B2["backup-from-snap2.cfg<br/>(offsets + checksums)"]
        B3["backup-from-snap3.cfg"]
        BLK[("blocks/*.blk<br/>2 MiB, compressed,<br/>checksum-addressed")]
    end
    S2 -->|"flatten snap1+snap2"| B2
    S3 -->|"only changed 2 MiB blocks"| B3
    B2 --> BLK
    B3 --> BLK
  • Backup targets: S3-compatible object storage (AWS S3, MinIO, GCS through its S3 API), NFS, SMB/CIFS and Azure Blob Storage. Since v1.8.0 several targets can be configured as BackupTarget CRs. Object storage is recommended over NFS, because there is nothing to mount and failover is simpler. URL formats are in the Reference.
  • Incremental backups: Each backup sends only 2 MiB blocks that changed since the previous backup. A block is re-sent whole even if only one 4 KiB page inside it changed. Blocks are addressed by checksum and shared across backups of the same volume. Deleting a backup leaves shared blocks, which a periodic garbage collection removes later. v1.10 made the block size configurable (2 or 16 MiB). A periodic full backup can be forced with full-backup-interval.
  • Restores produce a volume with a single snapshot. Backups do not carry snapshot history.
  • Recurring backups/snapshots: RecurringJob CRs are bound to volumes by group or label. The default group applies to volumes with no explicit job labels.

Encrypted Backups

Backups of encrypted volumes contain ciphertext. Restoring them requires the same encryption Secret in the target cluster.

Disaster Recovery Volumes

A DR (standby) volume in a second cluster follows a backup volume in the shared backupstore. It restores each new backup incrementally as the backupstore poll finds it. It cannot be snapshotted, backed up or bound to a PVC until it is activated. Activation fails if the last backup is not yet restored. After activation it is an ordinary volume and cannot go back to standby.

  • RPO is set by the recurring backup schedule of the source volume, for example hourly backups give a 1 h RPO.
  • RTO is driven by the backupstore poll interval. A long interval leaves more backups to replay at activation time.

System Backup

A Longhorn system backup (v1.4+) uploads a bundle of Longhorn-owned Kubernetes resources to the default backup target. The bundle holds Volumes, PVs, PVCs, StorageClasses, Settings, RecurringJobs, EngineImages, BackingImages, CRDs, RBAC objects, Deployments, DaemonSets and Services. It can also trigger volume backups (volumeBackupPolicy: if-not-present, always, disabled). A system restore rebuilds the Longhorn configuration in a fresh cluster and restores volumes from their latest backups. Longhorn Node CRs are not included, because each manager recreates its own. Upstream recommends a system backup before every upgrade.

Storage Sharding (Experimental, v1.12.1)

v1.12.1 adds an experimental sharded data layout on the V2 engine. Instead of full copies, written data is erasure-coded into data and parity chunks spread across nodes. A volume can then outgrow a single disk and use less raw capacity for the same fault tolerance. It is for evaluation only. It does not support backup and restore, cloning, DR volumes, live migration or replica scheduling, and expansion is limited to 10x the creation size. The new shardgroups and shards CRDs back it.

Upgrade Model

  • Two-step upgrade: first longhorn-manager (Helm or manifest), then the engine image per volume. Engine upgrade can be automatic (Concurrent Automatic Engine Upgrade Per Node Limit) or manual.
  • Live engine upgrade (V1 only): a new engine process starts beside the old one and takes over the replicas without detaching the volume.
  • One minor at a time: since v1.5.0 the manager blocks skipped minors (1.10 to 1.12) and downgrades. Plan for one upgrade per minor release.
  • V2 volumes must be detached for any upgrade in the 1.12 line.
  • 1.12.1 specifics: internal NetworkPolicies are created by default and can block traffic on enforcing CNIs. Legacy V2 linked clones become read-only. V2 volumes with backing images must be migrated before upgrading.

Key Architectural Properties

  • One engine per volume: Failure domains are isolated. A controller crash affects only one volume. The Instance Manager pod is the wider, per-node failure domain.
  • Microservices-based: Engine, replicas, Instance Manager and Share Manager are all orchestrated as Kubernetes resources and CRs.
  • Thin provisioning: Volumes consume only the space actually written. Over-provisioning is governed by Storage Over Provisioning Percentage (default 100) and Storage Minimal Available Percentage (default 25).
  • Crash-consistent snapshots and backups.
  • Live upgrades for V1 engines. V2 requires detach.
  • Kubernetes-only: Longhorn needs Kubernetes, privileged root pods and host packages. It is not a standalone SAN.

Performance and Scale Characteristics

  • Writes cost replicas x network. Each write crosses the network N-1 times and waits for the slowest replica. That is why upstream recommends 10 Gbps networking, a dedicated storage network, and 2 replicas for data-intensive apps. Use strict-local (1 replica, V1 only) for apps that replicate themselves, such as distributed databases.
  • Latency matters more than IOPS for stability. HDDs work but can destabilise volumes under rebuild or backup load.
  • Control-plane scale is bounded by attach/detach throughput and instance-manager resources, not by etcd. In the official v1.6 scale test, 3/6/9 worker clusters handled 699/1,114/1,561 volumes before pod start times crossed 4 minutes. The API server and etcd were not overloaded. Instance-manager pods, then manager pods, were the heaviest consumers. There is no published hard node limit. The "not for 100+ nodes" rule of thumb in older notes is not an upstream statement.
  • Maximum volume size is effectively bounded by rebuild time (24 h cap), not by a format limit.

Numbers and test conditions are in Reference: Benchmarks and Scalability.

Security Model

Longhorn provides volume-level encryption using Linux dm-crypt, integrates with Kubernetes RBAC, and since v1.12.1 ships internal NetworkPolicies and full mTLS coverage for instance-manager gRPC. As a Kubernetes-native system, Longhorn inherits much of its security posture from the cluster itself.

Volume Encryption (LUKS/dm-crypt)

Longhorn supports per-volume encryption for both Filesystem and Block volume modes, on V1 and V2 and for RWX volumes. It protects against physical disk theft and unauthorised access to replica data.

Longhorn volume encryption relies on three Linux/Kubernetes components:

  1. dm-crypt (Linux kernel module): Creates and manages encrypted block devices.
  2. cryptsetup (CLI utility): Formats and opens encrypted devices using the LUKS2 header.
  3. Kubernetes Secrets: Store encryption keys and are referenced by the StorageClass through CSI secret parameters.

When a volume is provisioned with encryption enabled:

  1. The CSI driver reads the encryption Secret (passphrase and cipher configuration) referenced by the StorageClass. A PVC stays Pending until that Secret exists.
  2. On node staging, cryptsetup opens the underlying Longhorn block device using the Secret.
  3. The resulting decrypted device is formatted and mounted into the Pod.
  4. On detach, the encrypted device is closed. Only ciphertext remains on disk, in replicas and in backups.

Encryption can be global (one Secret for every volume of a StorageClass) or per volume. Per-volume encryption uses the CSI template variables ${pvc.name} and ${pvc.namespace}, which the external-provisioner expands at provisioning time. Recipes are in How-to Guides. Secret keys and defaults (for example aes-xts-plain64, argon2i) are in the Reference.

Online Expansion of Encrypted Volumes

Online expansion needs the node-expand-secret-name / node-expand-secret-namespace StorageClass parameters. On Kubernetes v1.25 to v1.28 it also needs the CSINodeExpandSecret feature gate. It is GA in Kubernetes v1.29+.

Longhorn can also encrypt backing images (V1 only). The encrypted backing image carries a LUKS header, and volumes created from it are encrypted too.

CSI Driver Security

  • Sidecar containers: External provisioner, attacher, resizer and snapshotter sidecars handle CSI RPCs. They talk to the plugin over a Unix socket (/var/lib/kubelet/plugins/driver.longhorn.io/csi.sock). The plugin talks to the Longhorn Manager API.
  • Secret handling: Encryption keys and credentials are passed through the CSI secret mechanism, not written into CR status fields.
  • Node operations: The CSI node plugin needs privileged access to the host (/dev, /sys, /lib/modules, kubelet directories) for device staging and cryptsetup.

Network Exposure

  • The manager API (9500), the webhooks (9501, 9502) and the recovery backend (9503) are in-cluster services. The UI has no authentication of its own, so expose it only behind an authenticating ingress.
  • v1.12.1 creates ingress NetworkPolicies for manager, instance-manager, webhook, recovery-backend and backing-image components by default (networkPolicies.restrictInternalTraffic). They only take effect on a CNI that enforces NetworkPolicy. They can block Prometheus scrapes from other namespaces or unusual API-server source addresses, so check them after upgrading.
  • Replica traffic (10000-30000/TCP) is unencrypted block I/O. Upstream advises a dedicated storage network. When the longhorn-grpc-tls secret is set, control gRPC to instance managers uses mTLS, but replica data is still plaintext.
  • For the V2 Data Engine, SPDK needs huge pages and VFIO/UIO device access (vfio_pci, uio_pci_generic). Handing whole NVMe devices to a privileged user-space process increases the host attack surface.

Backup Target Credentials

  • S3: Access key and secret key in a Kubernetes Secret in longhorn-system, or IAM roles through kube2iam/kiam (AWS_IAM_ROLE_ARN). A custom CA is set with AWS_CERT.
  • NFS: No built-in authentication beyond export rules. Use network-level controls.
  • CIFS/SMB: Username and password in a Secret.
  • Azure Blob: Storage account name and key in a Secret.

Threat Model Summary

Threat Mitigation
Data exposure from stolen disks Volume-level LUKS2 encryption
Unauthorized API access Kubernetes RBAC, namespace isolation, internal NetworkPolicies (v1.12.1)
Unauthenticated UI access Authenticating ingress. Keep the longhorn-frontend Service internal
Network interception of replica traffic Dedicated storage network. Encrypt sensitive volumes (ciphertext on the wire)
Rogue client calling instance-manager gRPC longhorn-grpc-tls mTLS (full coverage in v1.12.1), NetworkPolicies
Webhook abuse from inside the cluster kubeAPIServerSourceCIDRs set to exact API server addresses
Compromised backup target Encrypted volumes give encrypted backups. Least-privilege S3 bucket policies
Privilege escalation via CSI driver Pod Security Admission on user namespaces. Restricted ServiceAccount RBAC
Host-level attack via V2 SPDK VFIO device isolation. Huge page reservation limits

The hardening checklist and RBAC table are in Reference: Hardening Checklist.

Sources