Skip to content

Flannel Explanation

What this page covers

How Flannel works and why it is built the way it is: the flanneld agent, the two subnet managers (Kubernetes API and etcd), the flannel CNI plugin, backend internals (VXLAN, host-gw, WireGuard, UDP, IPIP, IPsec), traffic management with iptables or nftables, dual-stack design, why Flannel has no built-in policy engine, and its security model. Look-up tables (config keys, flags, ports) live in Reference; tasks live in How-to Guides.

Overview

Flannel is a layer 3 network fabric for Kubernetes. CoreOS created it (the repository was originally coreos/flannel, and the flannel.alpha.coreos.com annotation prefix and /coreos.com/network etcd prefix are leftovers of that history). Today the flannel-io GitHub organization hosts it, and four maintainers affiliated with SUSE run it (GOVERNANCE.md).

Flannel does one job. It gives every node a pod subnet out of a larger cluster CIDR and moves packets between those subnets. The upstream README puts it plainly: Flannel "does not control how containers are networked to the host, only how the traffic is transported between hosts". It does not replace kube-proxy, and it has no eBPF dataplane. The flanneld binary also does not enforce NetworkPolicy. The project covers policy with an optional add-on, described in Network Policy Design.

Most people meet Flannel through K3s, which embeds Flannel in its binary as the default CNI. Canal (Calico policy plus Flannel networking) also uses it. RKE2 uses Canal by default and offers standalone Flannel as an option (integrations.md).

Component Diagram

This diagram shows the components on two nodes that use the kube subnet manager and the VXLAN backend, the default in the upstream manifest.

graph TB
    subgraph CP["Control plane"]
        KCM["kube-controller-manager<br/>(--allocate-node-cidrs)"]
        API["kube-apiserver<br/>(Node objects + annotations)"]
        CM["ConfigMap kube-flannel-cfg<br/>(net-conf.json, cni-conf.json)"]
    end

    subgraph N1["Node 1 (kube-flannel-ds pod, hostNetwork)"]
        F1["flanneld<br/>--kube-subnet-mgr --ip-masq"]
        ENV1["/run/flannel/subnet.env"]
        CNI1["flannel CNI plugin<br/>(/opt/cni/bin/flannel)"]
        BR1["bridge plugin -> cni0"]
        V1["flannel.1<br/>(VXLAN VTEP)"]
        NAT1["iptables or nftables<br/>(masquerade + forward)"]
    end

    subgraph N2["Node 2"]
        F2["flanneld"]
        V2["flannel.1"]
        BR2["cni0"]
    end

    KCM -->|"sets spec.podCIDR"| API
    CM -->|"mounted at /etc/kube-flannel/"| F1
    F1 <-->|"watch Nodes, write<br/>backend-data annotations"| API
    F2 <-->|"watch Nodes"| API
    F1 --> ENV1
    ENV1 --> CNI1
    CNI1 -->|"delegates"| BR1
    F1 -->|"routes, ARP, FDB"| V1
    F1 --> NAT1
    V1 <-->|"VXLAN over UDP 8472"| V2
    V2 --> BR2
    F2 -->|"routes, ARP, FDB"| V2

Core Components

flanneld

flanneld is a single Go binary that runs on every node. On Kubernetes it runs as the kube-flannel container of the kube-flannel-ds DaemonSet in the kube-flannel namespace, with hostNetwork: true and the NET_ADMIN and NET_RAW capabilities. On startup it:

  1. Picks the external interface (--iface, --iface-regex, --iface-can-reach, or the default-route interface) and the node's public IP. Leases are tied to that IP, so a changed IP needs a flanneld restart (troubleshooting.md).
  2. Reads the network config (net-conf.json in kube mode, /coreos.com/network/config in etcd mode).
  3. Acquires the node subnet from the subnet manager.
  4. Initialises the backend: creates the tunnel device (flannel.1, flannel-wg, flannel0, flannel.ipip) or nothing for host-gw.
  5. Installs masquerade and forward rules through the traffic manager (iptables by default, nftables when EnableNFTables is true).
  6. Writes /run/flannel/subnet.env. From v0.28.9, /readyz only returns 200 after the rules are installed and this file is written.
  7. Watches for other nodes' leases and programs routes, neighbour and FDB entries as nodes come and go.

Subnet Managers

Flannel has two ways to store the network config and per-node subnets. They behave differently, and older notes often mix them up.

Aspect Kube subnet manager (--kube-subnet-mgr) etcd subnet manager (default when the flag is absent)
Who allocates the node subnet kube-controller-manager (--allocate-node-cidrs --cluster-cidr) or kubeadm --pod-network-cidr. Flannel reads node.spec.podCIDR. flanneld itself, from Network/SubnetLen/SubnetMin/SubnetMax
Where the config lives /etc/kube-flannel/net-conf.json (from the kube-flannel-cfg ConfigMap) /coreos.com/network/config (prefix set by --etcd-prefix)
Where peer data lives Node annotations (flannel.alpha.coreos.com/backend-data, backend-type, public-ip, kube-subnet-manager, and others) Lease keys under /coreos.com/network/subnets/
Lease lifetime Tied to the Node object; no TTL renewal loop 24 h TTL, renewed --subnet-lease-renew-margin (60 min default) before expiry
Typical use Every Kubernetes install, including K3s Flannel outside Kubernetes (Docker hosts, legacy setups)

The README recommends the kube subnet manager because it removes the need for a separate etcd cluster. A frequent failure in kube mode is node <NODE_NAME> pod cidr not assigned. The node has no podCIDR because nothing on the control plane allocates one.

Flannel CNI Plugin

Flannel ships its own small CNI plugin (flannel, image ghcr.io/flannel-io/flannel-cni-plugin, currently v1.9.1-flannel3). The install-cni-plugin init container copies it to /opt/cni/bin/flannel. The install-cni init container writes /etc/cni/net.d/10-flannel.conflist, which chains flannel with portmap.

When the kubelet calls the plugin for a new pod, the plugin:

  1. Reads /run/flannel/subnet.env (FLANNEL_NETWORK, FLANNEL_SUBNET, FLANNEL_MTU, FLANNEL_IPMASQ, plus the IPv6 variants).
  2. Builds a config for the standard bridge plugin with host-local IPAM over the node subnet. The upstream conflist sets hairpinMode: true and isDefaultGateway: true.
  3. Delegates to bridge, which creates the cni0 bridge and a veth pair and assigns the pod IP.

So "Flannel uses the bridge plugin" is only half the story: Flannel's plugin is a thin shim that turns subnet.env into a bridge plus host-local config. The standard CNI plugins (bridge, host-local, portmap) must be present in /opt/cni/bin.

Packet Flow (VXLAN)

This sequence shows a packet going from a pod on Node 1 to a pod on Node 2 with the VXLAN backend.

sequenceDiagram
    participant PodA as Pod A 10.244.0.2 (Node 1)
    participant BR1 as cni0 bridge (Node 1)
    participant K1 as Node 1 kernel (route, ARP, FDB)
    participant VT1 as flannel.1 VTEP (Node 1)
    participant NET as Underlay network
    participant VT2 as flannel.1 VTEP (Node 2)
    participant PodB as Pod B 10.244.1.3 (Node 2)

    PodA->>BR1: IP packet 10.244.0.2 to 10.244.1.3
    BR1->>K1: Route lookup finds 10.244.1.0/24 via 10.244.1.0 dev flannel.1 onlink
    K1->>K1: Static ARP entry maps 10.244.1.0 to the Node 2 VTEP MAC
    K1->>K1: FDB entry maps the VTEP MAC to the Node 2 public IP
    K1->>VT1: Hand frame to the VXLAN device
    VT1->>NET: Outer IP/UDP to Node 2 port 8472, VNI 1
    NET->>VT2: UDP datagram arrives on port 8472
    VT2->>VT2: Decapsulate, check VNI
    VT2->>PodB: Inner packet routed to cni0 and the pod veth

Data Flow Summary

Scenario Path
Pod-to-pod on the same node veth -> cni0 bridge -> veth (no encapsulation)
Pod-to-pod across nodes (VXLAN) veth -> cni0 -> route -> flannel.1 (encap) -> NIC -> flannel.1 (decap) -> cni0 -> veth
Pod-to-pod across nodes (host-gw) veth -> cni0 -> route via peer node IP -> NIC -> peer route -> cni0 -> veth
Pod-to-pod across nodes (WireGuard) veth -> cni0 -> route -> flannel-wg (encrypt) -> NIC -> flannel-wg (decrypt) -> cni0 -> veth
Pod-to-external veth -> cni0 -> MASQUERADE to node IP (when --ip-masq) -> NIC

VXLAN Internals

The comments in pkg/backend/vxlan/vxlan.go describe three generations of the design:

  1. v1: flanneld handled kernel L2-miss (ARP) and L3-miss (FDB) netlink callouts at packet time.
  2. v2: the L3-miss callout was removed. flanneld pre-populated FDB entries when it discovered a host.
  3. Current: no callouts at all. For each remote host, flanneld installs one route, one static ARP entry and one FDB entry through netlink. Table size grows linearly with the number of nodes, and the dataplane keeps working even while flanneld restarts.

Consequences worth knowing:

  • Route: 10.244.1.0/24 via 10.244.1.0 dev flannel.1 onlink. The gateway is the remote subnet's network address, which is the IP of the remote flannel.1. onlink tells the kernel to accept a next hop that is not on a connected subnet.
  • MAC learning is off by default (Learning: false). Entries are static and owned by flanneld.
  • Port 8472 is the Linux kernel's legacy VXLAN default, not IANA's 4789. Windows must use 4789 and a VNI of 4096 or higher.
  • MTU overhead is 50 bytes (encapOverhead = 50), so a 1500-byte NIC gives FLANNEL_MTU=1450.
  • DirectRouting installs plain host-gw style routes for peers on the same L2 subnet and uses VXLAN only for peers across routers.
  • Behind NAT, VXLAN UDP checksums can be corrupted. The documented workaround is ethtool -K flannel.1 tx-checksum-ip-generic off.

Backend Internals

A backend is chosen once, in Backend.Type, and "should not be changed at runtime" (backends.md). Upstream sorts backends into two groups:

  • Recommended: vxlan, host-gw, wireguard, and udp (UDP is for debugging only).
  • Experimental and unsupported: alloc, tencent-vpc, ipip, ipsec, and extension.

Older cloud backends (aws-vpc, gce, ali-vpc) were removed in v0.20.0 (2022-10): commit bb65738 "Remove obsolete cloud backends" (PR #1625) is first contained in the v0.20.0 tag.

host-gw

With host-gw, flanneld installs remote-subnet via remote-node-IP routes on the physical interface. There is no encapsulation, no MTU penalty, and no tunnel device. All nodes therefore need direct L2 adjacency, because the next hop must be on-link. Upstream notes that host-gw "typically can't be used in cloud environments". Cloud VPCs drop packets whose source or destination IPs they do not know unless you disable source/destination checks or program VPC routes.

WireGuard

The WireGuard backend creates in-kernel WireGuard interfaces: flannel-wg, plus flannel-wg-v6 in the default separate mode. It publishes each node's public key in the backend-data annotation. It generates a private key into /run/flannel/wgkey (override with WIREGUARD_KEY_FILE) and adds a peer per remote node, with that node's pod subnet as AllowedIPs. The listen ports are 51820 for IPv4 and 51821 for IPv6. The MTU overhead constant is 80 bytes. Kernels older than 5.6 need the out-of-tree WireGuard module. Mode chooses between separate v4 and v6 tunnels (separate) and a single tunnel (auto, ipv4, or ipv6). A PSK adds a symmetric pre-shared key on top of the Curve25519 handshake.

UDP

The UDP backend is the original 2014 design. It uses a TUN device (flannel0) and a userspace C proxy that wraps packets in UDP port 8285, with 28 bytes of overhead. Every packet crosses the kernel/userspace boundary twice, so it is the slowest backend. Upstream keeps it "for debugging only or for very old kernels that don't support VXLAN". It is not formally deprecated, but on any architecture other than amd64 and on Windows it returns UDP backend is not supported on this architecture.

IPIP, IPsec, alloc, extension

  • IPIP (flannel.ipip, 20-byte overhead) is lighter than VXLAN but carries only IPv4 unicast. It supports DirectRouting. The kernel's tunl0 device also appears; this is expected.
  • IPsec uses in-kernel XFRM with strongSwan (charon) as the IKEv2 daemon and a PSK of at least 96 characters. It needs ESP (IP protocol 50), UDP 500 and UDP 4500. It is experimental. K3s has deprecated its ipsec backend option in favour of wireguard-native.
  • alloc only allocates subnets and forwards nothing. It is useful when something else programs the routes.
  • extension runs user shell commands on subnet add and remove. It is meant for prototyping. CVE-2026-32241 (High, CVSS 7.5) showed that before v0.28.2 a user allowed to set the backend-data Node annotation could inject commands that ran as root on every node using this backend. Other backends were not affected.

Subnet Allocation and Lease Lifecycle

This diagram shows how the cluster CIDR is split into node subnets and pod IPs with the default SubnetLen of 24.

graph TD
    CIDR["Network 10.244.0.0/16<br/>(net-conf.json)"]
    N1["Node 1: 10.244.0.0/24<br/>cni0 = 10.244.0.1"]
    N2["Node 2: 10.244.1.0/24<br/>cni0 = 10.244.1.1"]
    NN["Node N: 10.244.N.0/24"]
    CIDR --> N1
    CIDR --> N2
    CIDR --> NN
    N1 -->|"host-local IPAM<br/>10.244.0.2 - .254"| P1["Pods on Node 1"]
    N2 -->|"host-local IPAM<br/>10.244.1.2 - .254"| P2["Pods on Node 2"]

SubnetLen defaults to 24 unless Network is smaller than a /22, in which case it is two bits longer than the network. The IPv6 default is /64. In kube mode the node subnet size comes from kube-controller-manager's --node-cidr-mask-size, and Flannel just uses whatever podCIDR the node carries.

In etcd mode, flanneld owns the lease lifecycle, as this state diagram shows.

stateDiagram-v2
    [*] --> Acquiring: flanneld starts (etcd mode)
    Acquiring --> Active: lease written with 24h TTL
    Acquiring --> Active: previous lease for same public IP reused
    Active --> Active: renew within subnet-lease-renew-margin (60 min default)
    Active --> Expired: node down, TTL elapses
    Expired --> [*]: key deleted by etcd, subnet free for reallocation
    Expired --> Acquiring: node returns

Traffic Management: iptables and nftables

flanneld installs two kinds of rules:

  • Masquerade (with --ip-masq): traffic from pods to destinations outside the flannel network is SNATed to the node IP. Traffic inside the overlay is not. By default MASQUERADE uses --random-fully; turn this off with --ip-masq-fully-random-disable.
  • Forward accept (--iptables-forward-rules, default true): accepts forwarding to and from the pod network. This matters on hosts whose FORWARD policy is DROP, which Docker has set since v1.13.

With iptables, the rules live in the FLANNEL-POSTRTG and FLANNEL-FWD chains and are resynced every --iptables-resync seconds (5 by default). The nftables ADR gives the reasons for adding nftables: iptables rules must be purged and re-added to keep their order, they interfere with other iptables users such as kube-proxy and kube-router, and distributions are moving to nftables. Since v0.25.0 (2024-04-08), "EnableNFTables": true makes Flannel program its own nftables tables (flannel-ipv4 and flannel-ipv6, managed through the knftables library) instead. The feature is still marked EXPERIMENTAL and iptables remains the default. It pairs naturally with kube-proxy's nftables mode, which is GA since Kubernetes v1.33.

Dual-Stack and IPv6

Dual-stack came later and is layered on top of the same model:

  • EnableIPv6 with IPv6Network allocates a second subnet per node. Setting EnableIPv4: false gives IPv6-only.
  • Only vxlan, wireguard and host-gw (Linux) support dual-stack. VXLAN creates a second device, flannel-v6.<VNI>. WireGuard's default separate mode creates flannel-wg-v6 on port 51821.
  • Nodes need an IPv4 and an IPv6 address and default route on the main interface.
  • With public (GUA) IPv6 pod ranges, Flannel does not advertise routes to the outside. Upstream routers must route IPv6Network to the cluster, or you enable IPv6 masquerade (--flannel-ipv6-masq on K3s).

Network Policy Design

Flannel deliberately keeps policy out of flanneld. Policy is filtering, not connectivity, and the project treats it as a separate, composable controller. The diagram below shows the four supported ways to get NetworkPolicy enforcement on a Flannel network.

flowchart LR
    FL["Flannel<br/>(connectivity only)"]
    KNP["kube-network-policies<br/>(Helm netpol.enabled, since v0.25.5)"]
    KR["kube-router netpol library<br/>(embedded in K3s)"]
    CAN["Calico Felix<br/>(Canal)"]
    CIL["Cilium<br/>(CNI chaining)"]
    FL --> KNP
    FL --> KR
    FL --> CAN
    FL --> CIL
  • kube-network-policies (Kubernetes SIG Network): the Flannel Helm chart's netpol.enabled=true adds this controller as an extra container (image registry.k8s.io/networking/kube-network-policies, v1.0.0 in the current chart). Available since Flannel v0.25.5.
  • K3s embedded controller: K3s runs kube-router's netpol controller library (only that part of kube-router) next to its embedded Flannel. Disable it with --disable-network-policy.
  • Canal: Calico's Felix enforces policy and Flannel carries inter-node traffic. It is RKE2's default.
  • Cilium chaining: Cilium runs in chaining mode on top of Flannel for policy and visibility.

Without one of these, Kubernetes accepts NetworkPolicy objects and nothing enforces them. This is the most common security surprise on vanilla Flannel.

Security Model

Flannel's threat model is simple: the overlay is flat and trusted.

Threat Default exposure Mitigation
Pod-to-pod lateral movement Every pod can reach every pod in every namespace Add a policy controller (see above)
Sniffing inter-node traffic VXLAN, host-gw, UDP and IPIP send cleartext on the underlay wireguard backend (or experimental ipsec), or an encrypted underlay
Packet injection into the overlay Anyone who can reach UDP 8472 on a node can inject VXLAN frames. K3s docs warn not to expose the port publicly. Firewall 8472/51820/51821 to node IPs only
ARP spoofing on cni0 Pods with NET_RAW can spoof ARP on the node's L2 bridge Pod Security restricted profile drops NET_RAW
Tampering with Node annotations Annotations carry peer endpoints and keys. With extension, they carried commands (CVE-2026-32241). Keep flannel at the latest release; restrict nodes/patch RBAC
Privileged DaemonSet kube-flannel needs NET_ADMIN, NET_RAW, hostNetwork and hostPath mounts, so the namespace must be Pod Security privileged Restrict who can create pods in kube-flannel
etcd tampering (etcd mode) Whoever can write /coreos.com/network/ can hijack subnets etcd mTLS (--etcd-certfile/keyfile/cafile) and scoped etcd users

WireGuard in Flannel encrypts only traffic between nodes. Same-node pod traffic crosses cni0 in cleartext, and there is no per-namespace or identity-aware encryption. Encryption is comparable to Calico's or Cilium's WireGuard modes, but those projects add identity and policy on top.

Upstream SECURITY.md patches only the latest release. That makes "stay on the latest patch" the only supported security posture.

What Flannel Does Not Provide

Feature gaps by design

  • NetworkPolicy in flanneld itself. Add a controller as shown above.
  • eBPF dataplane or kube-proxy replacement. Services still go through kube-proxy (iptables, IPVS or nftables mode).
  • Observability. No flow logs or Hubble-style service map. Only logs (-v) and the /healthz and /readyz endpoints.
  • BGP or route advertisement. Pod CIDRs are not advertised to the physical network.
  • Multi-cluster. No ClusterMesh equivalent.
  • Flexible IPAM. One contiguous subnet per node, with no IP pools, block borrowing or per-namespace pools like Calico IPAM.

Performance Characteristics

Upstream gives only qualitative guidance (troubleshooting.md):

  • Backend choice matters most. "If encapsulation is used, vxlan will always perform better than udp. For maximum data plane performance, avoid encapsulation" (host-gw, or VXLAN with DirectRouting).
  • MTU matters next. Check the NIC MTU, then FLANNEL_MTU in subnet.env, then the pod veth MTU. Jumbo frames help encapsulated backends most.
  • Control-plane load is light. Flannel "is known to scale to a very large number of hosts". A slow start of pods on a new node usually points to the datastore (API server or etcd) rather than to Flannel. For clusters with more nodes than EVENT_QUEUE_DEPTH (5000 by default), raise that value or set CONT_WHEN_CACHE_NOT_READY=true.

Upstream's ROADMAP lists "a public benchmark suite for backend comparison" as a future item. The project publishes no official numbers yet (checked 2026-09-27); Reference > Performance Estimates keeps only the qualitative ordering and the sourced resource requests.

Design Trade-offs

Decision Benefit Cost
One subnet per node, host-local IPAM No IPAM coordination on the pod creation path Wastes addresses; node count is capped by Network / SubnetLen
Kernel dataplane (VXLAN, routes, WireGuard) instead of eBPF Works on nearly any kernel and architecture, including arm, riscv64 and s390x No L7, no eBPF service handling, no flow visibility
Static ARP and FDB entries No netlink callouts; the dataplane survives flanneld restarts O(nodes) entries per host
Policy left to other controllers Small, auditable core One more component to run for any multi-tenant cluster
Backend fixed at install Simple state model Changing backend means a disruptive reconfiguration

Sources