Flannel Explanation¶
What this page covers
How Flannel works and why it is built the way it is: the flanneld agent, the two subnet managers (Kubernetes API and etcd), the flannel CNI plugin, backend internals (VXLAN, host-gw, WireGuard, UDP, IPIP, IPsec), traffic management with iptables or nftables, dual-stack design, why Flannel has no built-in policy engine, and its security model. Look-up tables (config keys, flags, ports) live in Reference; tasks live in How-to Guides.
Overview¶
Flannel is a layer 3 network fabric for Kubernetes. CoreOS created it (the repository was originally coreos/flannel, and the flannel.alpha.coreos.com annotation prefix and /coreos.com/network etcd prefix are leftovers of that history). Today the flannel-io GitHub organization hosts it, and four maintainers affiliated with SUSE run it (GOVERNANCE.md).
Flannel does one job. It gives every node a pod subnet out of a larger cluster CIDR and moves packets between those subnets. The upstream README puts it plainly: Flannel "does not control how containers are networked to the host, only how the traffic is transported between hosts". It does not replace kube-proxy, and it has no eBPF dataplane. The flanneld binary also does not enforce NetworkPolicy. The project covers policy with an optional add-on, described in Network Policy Design.
Most people meet Flannel through K3s, which embeds Flannel in its binary as the default CNI. Canal (Calico policy plus Flannel networking) also uses it. RKE2 uses Canal by default and offers standalone Flannel as an option (integrations.md).
Component Diagram¶
This diagram shows the components on two nodes that use the kube subnet manager and the VXLAN backend, the default in the upstream manifest.
graph TB
subgraph CP["Control plane"]
KCM["kube-controller-manager<br/>(--allocate-node-cidrs)"]
API["kube-apiserver<br/>(Node objects + annotations)"]
CM["ConfigMap kube-flannel-cfg<br/>(net-conf.json, cni-conf.json)"]
end
subgraph N1["Node 1 (kube-flannel-ds pod, hostNetwork)"]
F1["flanneld<br/>--kube-subnet-mgr --ip-masq"]
ENV1["/run/flannel/subnet.env"]
CNI1["flannel CNI plugin<br/>(/opt/cni/bin/flannel)"]
BR1["bridge plugin -> cni0"]
V1["flannel.1<br/>(VXLAN VTEP)"]
NAT1["iptables or nftables<br/>(masquerade + forward)"]
end
subgraph N2["Node 2"]
F2["flanneld"]
V2["flannel.1"]
BR2["cni0"]
end
KCM -->|"sets spec.podCIDR"| API
CM -->|"mounted at /etc/kube-flannel/"| F1
F1 <-->|"watch Nodes, write<br/>backend-data annotations"| API
F2 <-->|"watch Nodes"| API
F1 --> ENV1
ENV1 --> CNI1
CNI1 -->|"delegates"| BR1
F1 -->|"routes, ARP, FDB"| V1
F1 --> NAT1
V1 <-->|"VXLAN over UDP 8472"| V2
V2 --> BR2
F2 -->|"routes, ARP, FDB"| V2
Core Components¶
flanneld¶
flanneld is a single Go binary that runs on every node. On Kubernetes it runs as the kube-flannel container of the kube-flannel-ds DaemonSet in the kube-flannel namespace, with hostNetwork: true and the NET_ADMIN and NET_RAW capabilities. On startup it:
- Picks the external interface (
--iface,--iface-regex,--iface-can-reach, or the default-route interface) and the node's public IP. Leases are tied to that IP, so a changed IP needs aflanneldrestart (troubleshooting.md). - Reads the network config (
net-conf.jsonin kube mode,/coreos.com/network/configin etcd mode). - Acquires the node subnet from the subnet manager.
- Initialises the backend: creates the tunnel device (
flannel.1,flannel-wg,flannel0,flannel.ipip) or nothing for host-gw. - Installs masquerade and forward rules through the traffic manager (iptables by default, nftables when
EnableNFTablesis true). - Writes
/run/flannel/subnet.env. From v0.28.9,/readyzonly returns 200 after the rules are installed and this file is written. - Watches for other nodes' leases and programs routes, neighbour and FDB entries as nodes come and go.
Subnet Managers¶
Flannel has two ways to store the network config and per-node subnets. They behave differently, and older notes often mix them up.
| Aspect | Kube subnet manager (--kube-subnet-mgr) |
etcd subnet manager (default when the flag is absent) |
|---|---|---|
| Who allocates the node subnet | kube-controller-manager (--allocate-node-cidrs --cluster-cidr) or kubeadm --pod-network-cidr. Flannel reads node.spec.podCIDR. |
flanneld itself, from Network/SubnetLen/SubnetMin/SubnetMax |
| Where the config lives | /etc/kube-flannel/net-conf.json (from the kube-flannel-cfg ConfigMap) |
/coreos.com/network/config (prefix set by --etcd-prefix) |
| Where peer data lives | Node annotations (flannel.alpha.coreos.com/backend-data, backend-type, public-ip, kube-subnet-manager, and others) |
Lease keys under /coreos.com/network/subnets/ |
| Lease lifetime | Tied to the Node object; no TTL renewal loop | 24 h TTL, renewed --subnet-lease-renew-margin (60 min default) before expiry |
| Typical use | Every Kubernetes install, including K3s | Flannel outside Kubernetes (Docker hosts, legacy setups) |
The README recommends the kube subnet manager because it removes the need for a separate etcd cluster. A frequent failure in kube mode is node <NODE_NAME> pod cidr not assigned. The node has no podCIDR because nothing on the control plane allocates one.
Flannel CNI Plugin¶
Flannel ships its own small CNI plugin (flannel, image ghcr.io/flannel-io/flannel-cni-plugin, currently v1.9.1-flannel3). The install-cni-plugin init container copies it to /opt/cni/bin/flannel. The install-cni init container writes /etc/cni/net.d/10-flannel.conflist, which chains flannel with portmap.
When the kubelet calls the plugin for a new pod, the plugin:
- Reads
/run/flannel/subnet.env(FLANNEL_NETWORK,FLANNEL_SUBNET,FLANNEL_MTU,FLANNEL_IPMASQ, plus the IPv6 variants). - Builds a config for the standard
bridgeplugin withhost-localIPAM over the node subnet. The upstream conflist setshairpinMode: trueandisDefaultGateway: true. - Delegates to
bridge, which creates thecni0bridge and a veth pair and assigns the pod IP.
So "Flannel uses the bridge plugin" is only half the story: Flannel's plugin is a thin shim that turns subnet.env into a bridge plus host-local config. The standard CNI plugins (bridge, host-local, portmap) must be present in /opt/cni/bin.
Packet Flow (VXLAN)¶
This sequence shows a packet going from a pod on Node 1 to a pod on Node 2 with the VXLAN backend.
sequenceDiagram
participant PodA as Pod A 10.244.0.2 (Node 1)
participant BR1 as cni0 bridge (Node 1)
participant K1 as Node 1 kernel (route, ARP, FDB)
participant VT1 as flannel.1 VTEP (Node 1)
participant NET as Underlay network
participant VT2 as flannel.1 VTEP (Node 2)
participant PodB as Pod B 10.244.1.3 (Node 2)
PodA->>BR1: IP packet 10.244.0.2 to 10.244.1.3
BR1->>K1: Route lookup finds 10.244.1.0/24 via 10.244.1.0 dev flannel.1 onlink
K1->>K1: Static ARP entry maps 10.244.1.0 to the Node 2 VTEP MAC
K1->>K1: FDB entry maps the VTEP MAC to the Node 2 public IP
K1->>VT1: Hand frame to the VXLAN device
VT1->>NET: Outer IP/UDP to Node 2 port 8472, VNI 1
NET->>VT2: UDP datagram arrives on port 8472
VT2->>VT2: Decapsulate, check VNI
VT2->>PodB: Inner packet routed to cni0 and the pod veth
Data Flow Summary¶
| Scenario | Path |
|---|---|
| Pod-to-pod on the same node | veth -> cni0 bridge -> veth (no encapsulation) |
| Pod-to-pod across nodes (VXLAN) | veth -> cni0 -> route -> flannel.1 (encap) -> NIC -> flannel.1 (decap) -> cni0 -> veth |
| Pod-to-pod across nodes (host-gw) | veth -> cni0 -> route via peer node IP -> NIC -> peer route -> cni0 -> veth |
| Pod-to-pod across nodes (WireGuard) | veth -> cni0 -> route -> flannel-wg (encrypt) -> NIC -> flannel-wg (decrypt) -> cni0 -> veth |
| Pod-to-external | veth -> cni0 -> MASQUERADE to node IP (when --ip-masq) -> NIC |
VXLAN Internals¶
The comments in pkg/backend/vxlan/vxlan.go describe three generations of the design:
- v1:
flanneldhandled kernel L2-miss (ARP) and L3-miss (FDB) netlink callouts at packet time. - v2: the L3-miss callout was removed.
flanneldpre-populated FDB entries when it discovered a host. - Current: no callouts at all. For each remote host,
flanneldinstalls one route, one static ARP entry and one FDB entry through netlink. Table size grows linearly with the number of nodes, and the dataplane keeps working even whileflanneldrestarts.
Consequences worth knowing:
- Route:
10.244.1.0/24 via 10.244.1.0 dev flannel.1 onlink. The gateway is the remote subnet's network address, which is the IP of the remoteflannel.1.onlinktells the kernel to accept a next hop that is not on a connected subnet. - MAC learning is off by default (
Learning: false). Entries are static and owned byflanneld. - Port 8472 is the Linux kernel's legacy VXLAN default, not IANA's 4789. Windows must use 4789 and a VNI of 4096 or higher.
- MTU overhead is 50 bytes (
encapOverhead = 50), so a 1500-byte NIC givesFLANNEL_MTU=1450. - DirectRouting installs plain host-gw style routes for peers on the same L2 subnet and uses VXLAN only for peers across routers.
- Behind NAT, VXLAN UDP checksums can be corrupted. The documented workaround is
ethtool -K flannel.1 tx-checksum-ip-generic off.
Backend Internals¶
A backend is chosen once, in Backend.Type, and "should not be changed at runtime" (backends.md). Upstream sorts backends into two groups:
- Recommended:
vxlan,host-gw,wireguard, andudp(UDP is for debugging only). - Experimental and unsupported:
alloc,tencent-vpc,ipip,ipsec, andextension.
Older cloud backends (aws-vpc, gce, ali-vpc) were removed in v0.20.0 (2022-10): commit bb65738 "Remove obsolete cloud backends" (PR #1625) is first contained in the v0.20.0 tag.
host-gw¶
With host-gw, flanneld installs remote-subnet via remote-node-IP routes on the physical interface. There is no encapsulation, no MTU penalty, and no tunnel device. All nodes therefore need direct L2 adjacency, because the next hop must be on-link. Upstream notes that host-gw "typically can't be used in cloud environments". Cloud VPCs drop packets whose source or destination IPs they do not know unless you disable source/destination checks or program VPC routes.
WireGuard¶
The WireGuard backend creates in-kernel WireGuard interfaces: flannel-wg, plus flannel-wg-v6 in the default separate mode. It publishes each node's public key in the backend-data annotation. It generates a private key into /run/flannel/wgkey (override with WIREGUARD_KEY_FILE) and adds a peer per remote node, with that node's pod subnet as AllowedIPs. The listen ports are 51820 for IPv4 and 51821 for IPv6. The MTU overhead constant is 80 bytes. Kernels older than 5.6 need the out-of-tree WireGuard module. Mode chooses between separate v4 and v6 tunnels (separate) and a single tunnel (auto, ipv4, or ipv6). A PSK adds a symmetric pre-shared key on top of the Curve25519 handshake.
UDP¶
The UDP backend is the original 2014 design. It uses a TUN device (flannel0) and a userspace C proxy that wraps packets in UDP port 8285, with 28 bytes of overhead. Every packet crosses the kernel/userspace boundary twice, so it is the slowest backend. Upstream keeps it "for debugging only or for very old kernels that don't support VXLAN". It is not formally deprecated, but on any architecture other than amd64 and on Windows it returns UDP backend is not supported on this architecture.
IPIP, IPsec, alloc, extension¶
- IPIP (
flannel.ipip, 20-byte overhead) is lighter than VXLAN but carries only IPv4 unicast. It supportsDirectRouting. The kernel'stunl0device also appears; this is expected. - IPsec uses in-kernel XFRM with strongSwan (charon) as the IKEv2 daemon and a PSK of at least 96 characters. It needs ESP (IP protocol 50), UDP 500 and UDP 4500. It is experimental. K3s has deprecated its
ipsecbackend option in favour ofwireguard-native. - alloc only allocates subnets and forwards nothing. It is useful when something else programs the routes.
- extension runs user shell commands on subnet add and remove. It is meant for prototyping. CVE-2026-32241 (High, CVSS 7.5) showed that before v0.28.2 a user allowed to set the
backend-dataNode annotation could inject commands that ran as root on every node using this backend. Other backends were not affected.
Subnet Allocation and Lease Lifecycle¶
This diagram shows how the cluster CIDR is split into node subnets and pod IPs with the default SubnetLen of 24.
graph TD
CIDR["Network 10.244.0.0/16<br/>(net-conf.json)"]
N1["Node 1: 10.244.0.0/24<br/>cni0 = 10.244.0.1"]
N2["Node 2: 10.244.1.0/24<br/>cni0 = 10.244.1.1"]
NN["Node N: 10.244.N.0/24"]
CIDR --> N1
CIDR --> N2
CIDR --> NN
N1 -->|"host-local IPAM<br/>10.244.0.2 - .254"| P1["Pods on Node 1"]
N2 -->|"host-local IPAM<br/>10.244.1.2 - .254"| P2["Pods on Node 2"]
SubnetLen defaults to 24 unless Network is smaller than a /22, in which case it is two bits longer than the network. The IPv6 default is /64. In kube mode the node subnet size comes from kube-controller-manager's --node-cidr-mask-size, and Flannel just uses whatever podCIDR the node carries.
In etcd mode, flanneld owns the lease lifecycle, as this state diagram shows.
stateDiagram-v2
[*] --> Acquiring: flanneld starts (etcd mode)
Acquiring --> Active: lease written with 24h TTL
Acquiring --> Active: previous lease for same public IP reused
Active --> Active: renew within subnet-lease-renew-margin (60 min default)
Active --> Expired: node down, TTL elapses
Expired --> [*]: key deleted by etcd, subnet free for reallocation
Expired --> Acquiring: node returns
Traffic Management: iptables and nftables¶
flanneld installs two kinds of rules:
- Masquerade (with
--ip-masq): traffic from pods to destinations outside the flannel network is SNATed to the node IP. Traffic inside the overlay is not. By default MASQUERADE uses--random-fully; turn this off with--ip-masq-fully-random-disable. - Forward accept (
--iptables-forward-rules, default true): accepts forwarding to and from the pod network. This matters on hosts whose FORWARD policy is DROP, which Docker has set since v1.13.
With iptables, the rules live in the FLANNEL-POSTRTG and FLANNEL-FWD chains and are resynced every --iptables-resync seconds (5 by default). The nftables ADR gives the reasons for adding nftables: iptables rules must be purged and re-added to keep their order, they interfere with other iptables users such as kube-proxy and kube-router, and distributions are moving to nftables. Since v0.25.0 (2024-04-08), "EnableNFTables": true makes Flannel program its own nftables tables (flannel-ipv4 and flannel-ipv6, managed through the knftables library) instead. The feature is still marked EXPERIMENTAL and iptables remains the default. It pairs naturally with kube-proxy's nftables mode, which is GA since Kubernetes v1.33.
Dual-Stack and IPv6¶
Dual-stack came later and is layered on top of the same model:
EnableIPv6withIPv6Networkallocates a second subnet per node. SettingEnableIPv4: falsegives IPv6-only.- Only
vxlan,wireguardandhost-gw(Linux) support dual-stack. VXLAN creates a second device,flannel-v6.<VNI>. WireGuard's defaultseparatemode createsflannel-wg-v6on port 51821. - Nodes need an IPv4 and an IPv6 address and default route on the main interface.
- With public (GUA) IPv6 pod ranges, Flannel does not advertise routes to the outside. Upstream routers must route
IPv6Networkto the cluster, or you enable IPv6 masquerade (--flannel-ipv6-masqon K3s).
Network Policy Design¶
Flannel deliberately keeps policy out of flanneld. Policy is filtering, not connectivity, and the project treats it as a separate, composable controller. The diagram below shows the four supported ways to get NetworkPolicy enforcement on a Flannel network.
flowchart LR
FL["Flannel<br/>(connectivity only)"]
KNP["kube-network-policies<br/>(Helm netpol.enabled, since v0.25.5)"]
KR["kube-router netpol library<br/>(embedded in K3s)"]
CAN["Calico Felix<br/>(Canal)"]
CIL["Cilium<br/>(CNI chaining)"]
FL --> KNP
FL --> KR
FL --> CAN
FL --> CIL
- kube-network-policies (Kubernetes SIG Network): the Flannel Helm chart's
netpol.enabled=trueadds this controller as an extra container (imageregistry.k8s.io/networking/kube-network-policies,v1.0.0in the current chart). Available since Flannel v0.25.5. - K3s embedded controller: K3s runs kube-router's netpol controller library (only that part of kube-router) next to its embedded Flannel. Disable it with
--disable-network-policy. - Canal: Calico's Felix enforces policy and Flannel carries inter-node traffic. It is RKE2's default.
- Cilium chaining: Cilium runs in chaining mode on top of Flannel for policy and visibility.
Without one of these, Kubernetes accepts NetworkPolicy objects and nothing enforces them. This is the most common security surprise on vanilla Flannel.
Security Model¶
Flannel's threat model is simple: the overlay is flat and trusted.
| Threat | Default exposure | Mitigation |
|---|---|---|
| Pod-to-pod lateral movement | Every pod can reach every pod in every namespace | Add a policy controller (see above) |
| Sniffing inter-node traffic | VXLAN, host-gw, UDP and IPIP send cleartext on the underlay | wireguard backend (or experimental ipsec), or an encrypted underlay |
| Packet injection into the overlay | Anyone who can reach UDP 8472 on a node can inject VXLAN frames. K3s docs warn not to expose the port publicly. | Firewall 8472/51820/51821 to node IPs only |
ARP spoofing on cni0 |
Pods with NET_RAW can spoof ARP on the node's L2 bridge |
Pod Security restricted profile drops NET_RAW |
| Tampering with Node annotations | Annotations carry peer endpoints and keys. With extension, they carried commands (CVE-2026-32241). |
Keep flannel at the latest release; restrict nodes/patch RBAC |
| Privileged DaemonSet | kube-flannel needs NET_ADMIN, NET_RAW, hostNetwork and hostPath mounts, so the namespace must be Pod Security privileged |
Restrict who can create pods in kube-flannel |
| etcd tampering (etcd mode) | Whoever can write /coreos.com/network/ can hijack subnets |
etcd mTLS (--etcd-certfile/keyfile/cafile) and scoped etcd users |
WireGuard in Flannel encrypts only traffic between nodes. Same-node pod traffic crosses cni0 in cleartext, and there is no per-namespace or identity-aware encryption. Encryption is comparable to Calico's or Cilium's WireGuard modes, but those projects add identity and policy on top.
Upstream SECURITY.md patches only the latest release. That makes "stay on the latest patch" the only supported security posture.
What Flannel Does Not Provide¶
Feature gaps by design
- NetworkPolicy in flanneld itself. Add a controller as shown above.
- eBPF dataplane or kube-proxy replacement. Services still go through kube-proxy (iptables, IPVS or nftables mode).
- Observability. No flow logs or Hubble-style service map. Only logs (
-v) and the/healthzand/readyzendpoints. - BGP or route advertisement. Pod CIDRs are not advertised to the physical network.
- Multi-cluster. No ClusterMesh equivalent.
- Flexible IPAM. One contiguous subnet per node, with no IP pools, block borrowing or per-namespace pools like Calico IPAM.
Performance Characteristics¶
Upstream gives only qualitative guidance (troubleshooting.md):
- Backend choice matters most. "If encapsulation is used,
vxlanwill always perform better thanudp. For maximum data plane performance, avoid encapsulation" (host-gw, or VXLAN withDirectRouting). - MTU matters next. Check the NIC MTU, then
FLANNEL_MTUinsubnet.env, then the pod veth MTU. Jumbo frames help encapsulated backends most. - Control-plane load is light. Flannel "is known to scale to a very large number of hosts". A slow start of pods on a new node usually points to the datastore (API server or etcd) rather than to Flannel. For clusters with more nodes than
EVENT_QUEUE_DEPTH(5000 by default), raise that value or setCONT_WHEN_CACHE_NOT_READY=true.
Upstream's ROADMAP lists "a public benchmark suite for backend comparison" as a future item. The project publishes no official numbers yet (checked 2026-09-27); Reference > Performance Estimates keeps only the qualitative ordering and the sourced resource requests.
Design Trade-offs¶
| Decision | Benefit | Cost |
|---|---|---|
| One subnet per node, host-local IPAM | No IPAM coordination on the pod creation path | Wastes addresses; node count is capped by Network / SubnetLen |
| Kernel dataplane (VXLAN, routes, WireGuard) instead of eBPF | Works on nearly any kernel and architecture, including arm, riscv64 and s390x | No L7, no eBPF service handling, no flow visibility |
| Static ARP and FDB entries | No netlink callouts; the dataplane survives flanneld restarts |
O(nodes) entries per host |
| Policy left to other controllers | Small, auditable core | One more component to run for any multi-tenant cluster |
| Backend fixed at install | Simple state model | Changing backend means a disruptive reconfiguration |
Sources¶
- Flannel README
- Backends
- Configuration
- Kubernetes integration
- Troubleshooting
- nftables ADR
- Network policy controller
- VXLAN backend source and design notes
- SECURITY.md and GHSA-vchx-5pr6-ffx2 / CVE-2026-32241
- K3s networking services (network policy controller)
- K3s requirements (ports, VXLAN exposure warning)