Explanation¶
What this page covers
How Cilium works and why it is built this way: agent, operator and CNI roles, the eBPF datapath (TC, XDP, socket hooks, veth vs netkit), identity-based policy, Hubble, Tetragon, Cluster Mesh, kube-proxy replacement, encryption, and the threat model. Look-up tables (ports, Helm values, map limits, IPAM and routing modes, benchmarks) are in Reference. Step-by-step tasks are in How-to Guides. Topic hub: Cilium.
Overview¶
Cilium is an eBPF-based CNI plugin for Kubernetes that provides networking, security, and observability. Traditional CNI plugins rely on iptables. Cilium instead runs eBPF programs inside the Linux kernel for packet forwarding, policy enforcement, load balancing, and tracing. It supports tunnel (VXLAN/Geneve) and native routing modes, and it can replace kube-proxy entirely.
The design has three main ideas:
- Programmable kernel datapath. The agent compiles and loads eBPF programs and keeps state in BPF maps. Lookups are hash-map based, not linear rule chains.
- Identity instead of IP. Pods get a numeric security identity derived from their labels. Policy is enforced on identities, so it survives pod churn and IP reuse.
- Visibility from the same hooks. The programs that forward and filter packets also emit flow events, which Hubble turns into flow logs, service maps and metrics.
Component Diagram¶
This diagram shows the Cilium 1.20 control plane and per-node components and how they relate to Kubernetes, Hubble, Cluster Mesh and Tetragon.
graph TB
subgraph K8S["Kubernetes control plane"]
KAPI["kube-apiserver<br/>(Pods, Services, EndpointSlices,<br/>CNP/CCNP, CiliumIdentity CRDs)"]
end
subgraph OPS["cilium-operator (Deployment, 2 replicas)"]
OPER["IPAM (cluster-pool, multi-pool, ENI, Azure),<br/>CRD registration, identity GC,<br/>Gateway API / Ingress controller"]
end
subgraph NODE["Each node"]
AGENT["cilium-agent (DaemonSet)<br/>compiles + loads eBPF,<br/>allocates identities, embeds Hubble server"]
ENVOY["cilium-envoy (DaemonSet)<br/>L7 policy, Ingress, Gateway API"]
CNI["cilium-cni binary<br/>(called by kubelet)"]
BPF["eBPF programs<br/>TC/tcx, XDP, cgroup socket hooks"]
MAPS["BPF maps<br/>ipcache, policy, lb, CT, NAT"]
end
subgraph OBS["Observability"]
RELAY["Hubble Relay :4245"]
UI["Hubble UI"]
end
subgraph CM["Cluster Mesh (optional)"]
CMAPI["clustermesh-apiserver<br/>etcd + apiserver + kvstoremesh"]
end
TETRA["Tetragon (separate DaemonSet)<br/>kprobes, tracepoints, LSM"]
KAPI -->|"watch"| AGENT
KAPI -->|"watch"| OPER
OPER -->|"pod CIDRs via CiliumNode"| AGENT
CNI -->|"create endpoint"| AGENT
AGENT -->|"load"| BPF
AGENT -->|"write"| MAPS
BPF -->|"read"| MAPS
BPF -->|"redirect L7 traffic"| ENVOY
AGENT -->|"flows :4244"| RELAY
RELAY --> UI
CMAPI -->|"remote cluster state"| AGENT
KAPI -->|"local state"| CMAPI
Core Components¶
Cilium Agent¶
The Cilium agent (cilium-agent) runs as a DaemonSet on every node. It is the per-node control plane:
- eBPF program management: compiles, loads, and attaches eBPF programs to TC (tcx on newer kernels), XDP, and cgroup socket hooks. It regenerates and swaps programs when endpoints or policies change.
- Endpoint tracking: watches the Kubernetes API for pods, creates Cilium endpoints, and assigns security identities.
- Identity allocation: pods with the same security-relevant labels share one numeric identity. Identities are stored as
CiliumIdentityCRDs (default) or in etcd (kvstore mode). - Service load balancing: programs BPF maps for ClusterIP, NodePort, LoadBalancer, and ExternalIP services. In kube-proxy replacement mode this fully replaces iptables service routing. The load-balancing control plane was redesigned in 1.18 to use less memory, and 1.20 reworked the backend representation so thousands of services can share a backend efficiently.
- Policy enforcement: turns Kubernetes NetworkPolicy, CiliumNetworkPolicy, CiliumClusterwideNetworkPolicy and (since 1.20) Kubernetes ClusterNetworkPolicy into per-endpoint BPF policy maps.
- Routing: programs kernel routes for native routing, or manages VXLAN/Geneve devices for overlay mode.
- Embedded services: the Hubble server (port 4244), the DNS proxy used for
toFQDNs, and the BGP control plane (GoBGP v4 since 1.20).
Cilium Operator¶
The Cilium Operator runs as a Deployment (2 replicas by default) and handles cluster-wide jobs:
- IPAM management: allocates pod CIDRs to nodes in cluster-pool and multi-pool mode, and cloud IPs/prefixes in ENI, Azure and AlibabaCloud modes.
- CRD lifecycle: registers and maintains
CiliumIdentity,CiliumEndpoint/CiliumEndpointSlice,CiliumNodeand other CRDs. - Garbage collection: removes stale identities and endpoints.
- Controllers: runs the Gateway API and Ingress controllers that turn Kubernetes resources into Envoy configuration, and LB-IPAM for LoadBalancer IPs.
CNI Plugin¶
The CNI binary (cilium-cni) is called by the kubelet (through the container runtime) when a pod is created. It:
- Calls the Cilium agent API to create an endpoint for the pod.
- Creates the pod device (a veth pair by default, or a netkit pair) and moves one end into the pod's network namespace.
- Waits for the agent to regenerate the endpoint's eBPF programs and policy so the pod starts with the correct rules.
Since 1.20 Cilium uses CNI spec version 1.0.0 by default.
cilium-envoy¶
L7 functions (HTTP/gRPC policy, TLS interception, Ingress, Gateway API, GAMMA) are handled by Envoy. For new installs Envoy runs as a separate cilium-envoy DaemonSet (envoy.enabled=true) rather than inside the agent pod, so it can be upgraded and resourced on its own. There is one Envoy per node, not a sidecar per pod. eBPF redirects only the flows that need L7 handling to it. Cilium 1.20 ships cilium-envoy v1.37.x and adds an ADS (Aggregated Discovery Service) server.
clustermesh-apiserver¶
With Cluster Mesh enabled, each cluster runs clustermesh-apiserver pods that expose local state to peer clusters and pull remote state in. See ClusterMesh.
eBPF Datapath¶
Cilium's datapath is built on eBPF. Programs are attached at several hook points in the kernel networking stack. The diagram shows where each hook sits on ingress and egress and which maps it consults.
graph TD
subgraph ING["Packet flow: ingress path"]
NIC["Network Interface<br/>(eth0)"]
XDP["XDP Hook<br/>(DDoS mitigation,<br/>early drop, LB acceleration)"]
TC_ING["TC ingress on host device<br/>(policy check,<br/>load balancing, routing)"]
L3["eBPF host routing<br/>(bpf_redirect_neigh / redirect_peer)"]
VETH["Pod device<br/>(veth or netkit)"]
end
subgraph EGR["Packet flow: egress path"]
POD_EG["Pod sends packet"]
TC_EG["TC hook on pod device<br/>(policy check, NAT, SNAT)"]
SOCK_LB["Socket-level LB<br/>cgroup connect/sendmsg hooks<br/>(connect-time LB, no per-packet NAT)"]
end
subgraph MAPS["BPF maps (shared state)"]
EP_MAP["Endpoint / ipcache maps<br/>IP -> identity"]
SVC_MAP["Service maps<br/>frontend -> backends"]
POL_MAP["Policy map<br/>identity + port -> verdict"]
CT_MAP["Connection tracking map"]
end
NIC -->|"raw packet"| XDP
XDP -->|"pass"| TC_ING
TC_ING -->|"lookup"| EP_MAP
TC_ING -->|"service resolve"| SVC_MAP
TC_ING -->|"policy check"| POL_MAP
TC_ING -->|"allowed"| L3
L3 -->|"local delivery"| VETH
POD_EG -->|"packet"| TC_EG
POD_EG -->|"connect()"| SOCK_LB
SOCK_LB -->|"direct to backend"| DEST["Destination Pod"]
TC_EG -->|"lookup"| CT_MAP
TC_EG -->|"policy check"| POL_MAP
TC_EG -->|"forward"| NIC
eBPF Hook Points¶
| Hook Point | Attachment | Purpose |
|---|---|---|
| XDP (eXpress Data Path) | Network driver, before skb allocation | Early drop, DDoS mitigation, XDP-accelerated NodePort/LoadBalancer (including DSR) |
| TC / tcx ingress | Physical, veth and netkit devices | Policy enforcement, service load balancing, routing decisions |
| TC / tcx egress | Physical and pod devices | SNAT, policy enforcement, tunnel encapsulation |
| cgroup socket hooks | cgroup v2 root | Connect-time load balancing (connect(), sendmsg()), which skips per-packet NAT |
| netkit peer programs | Inside the netkit device (1.20: bpf.datapathMode=netkit) |
Pod programs run in the device itself, removing the veth hop |
When netkit is enabled, Cilium uses tcx (BPF links) for attachments on all other devices too.
BPF Maps¶
BPF maps are the shared data structures between the Cilium agent (userspace) and the eBPF programs (kernel):
- Endpoint map: local pod IPs to endpoint ID, MAC and interface index.
- IP cache: every known IP (local and remote, across clusters) to its security identity and tunnel endpoint.
- Service maps: ClusterIP/NodePort/LoadBalancer frontends and their backends.
- Policy map (per endpoint): allowed identity + port + protocol entries. 1.20 added wildcard entries so the
world,remote-node,clusterand newcluster-meshentities use far fewer entries. - Connection tracking and NAT maps: stateful policy and NAT state.
- Tunnel map: remote node IPs to tunnel endpoints (VXLAN/Geneve mode).
BPF map memory
BPF maps are created with fixed upper limits and allocated in kernel memory. At large scale the CT, NAT and IP cache maps are the big consumers. Size them with bpf.mapDynamicSizeRatio (default 0.25% of node memory) or explicit values such as bpf.ctTcpMax, bpf.ctAnyMax, bpf.natMax and bpf.policyMapMax. Default limits are listed in Reference: eBPF Map Limits.
veth vs netkit¶
By default each pod gets a veth pair. A packet leaving the pod crosses the veth into the host namespace, where Cilium's TC program handles it. netkit (Linux 6.8+) is a device type built for Cilium: the pod's BPF program runs inside the netkit peer itself. With eBPF host routing this makes the namespace switch almost free, so pods get near host-namespace throughput and latency. It also supports BIG TCP.
Status in 1.20: beta. bpf.datapathMode accepts veth (default), netkit (L3), netkit-l2, and auto (new in 1.20: use netkit if the kernel supports it, otherwise veth). netkit requires eBPF host routing and does not work with bpf.tproxy. Existing pods cannot switch in place: veth and netkit cannot be mixed on one node. Enable it on new nodes (per-node config) or drain nodes first. The Cilium docs say veth mode will be deprecated once 6.8+ kernels are common.
How It Works¶
eBPF data plane internals, packet flow, the Hubble observability pipeline, and Tetragon enforcement.
eBPF Data Plane¶
Traditional CNIs use iptables, a linear chain of rules that gets slower as rules grow (O(n)). Cilium replaces this with eBPF hash maps that give constant-time O(1) lookups in kernel space. This diagram contrasts the two lookup models.
flowchart LR
subgraph Traditional["iptables-based CNI"]
PKT1["Packet"] --> R1["Rule 1"] --> R2["Rule 2"] --> R3["Rule 3"] --> RN["Rule N<br/>(O(n) traversal)"]
end
subgraph Cilium_DP["Cilium eBPF Data Plane"]
PKT2["Packet"] --> MAP["eBPF Hash Map<br/>(O(1) lookup)"] --> Action["Allow / Drop / Redirect"]
end
style Traditional fill:#c62828,color:#fff
style Cilium_DP fill:#2e7d32,color:#fff
Packet Flow — Pod-to-Pod (Same Node)¶
With eBPF host routing, a packet between two pods on the same node never goes through the host's iptables or routing stack. The sequence shows the veth case.
sequenceDiagram
participant PodA as Pod A
participant LXC_A as lxc device of Pod A (host side)
participant BPF as Cilium eBPF (from-container)
participant Maps as ipcache + policy maps
participant PodB as Pod B
PodA->>LXC_A: Send packet
LXC_A->>BPF: TC hook runs from-container program
BPF->>Maps: Resolve destination identity, check egress policy of A
BPF->>Maps: Check ingress policy of B, create CT entry
BPF->>PodB: bpf_redirect_peer into Pod B namespace
Note over BPF,PodB: Host network stack and iptables are bypassed
Packet Flow — Pod-to-Service (ClusterIP)¶
With socket-level load balancing (enabled by kube-proxy replacement), the service is resolved once at connect(). Packets then carry the backend IP from the start, so no per-packet DNAT is needed.
sequenceDiagram
participant Pod as Client Pod
participant Sock as cgroup connect hook
participant SVC as Service map
participant TC as TC eBPF (from-container)
participant Backend as Backend Pod
Pod->>Sock: connect() to ClusterIP:port
Sock->>SVC: Look up frontend, pick backend (random or Maglev)
Sock-->>Pod: Socket now bound to backend IP:port
Pod->>TC: Packets addressed to backend
TC->>Backend: Policy check, then forward directly
Note over Sock,SVC: Resolved once per connection, not per packet
Hubble Observability Pipeline¶
Datapath programs emit events to the agent through a perf event buffer (ring buffer support was added in 1.18). The Hubble server inside each agent turns them into flows. Relay aggregates flows cluster-wide. Metrics and flow export come from each agent.
flowchart TB
subgraph Kernel["Kernel space"]
eBPF_H["eBPF programs<br/>(TC, socket, XDP)"]
PerfBuf["Perf event buffer<br/>(trace, drop, policy verdict events)"]
end
subgraph Userspace["Hubble stack"]
HubbleAgent["Hubble server :4244<br/>(embedded in cilium-agent)"]
HubbleRelay["Hubble Relay :4245<br/>(cluster-wide aggregation)"]
HubbleUI["Hubble UI<br/>(service map, flow table)"]
CLI["hubble CLI"]
end
subgraph Export["Export"]
Prom["Prometheus<br/>(Hubble metrics per agent)"]
SIEM["Flow exporter<br/>(JSON file for log / SIEM pipelines)"]
end
eBPF_H -->|"perf events"| PerfBuf
PerfBuf --> HubbleAgent
HubbleAgent -->|"gRPC"| HubbleRelay
HubbleRelay --> HubbleUI
HubbleRelay --> CLI
HubbleAgent --> Prom
HubbleAgent --> SIEM
style Kernel fill:#f9a825,color:#000
style Userspace fill:#7b1fa2,color:#fff
Hubble¶
Hubble is Cilium's built-in network observability layer:
- Hubble server: embedded in the Cilium agent. Serves flow events (L3/L4/L7 metadata, policy verdicts) over gRPC on port 4244.
- Hubble Relay: a Deployment that connects to every node's Hubble server and serves a cluster-wide gRPC API on port 4245.
- Hubble UI: web interface with a service dependency map, flow table, and policy view.
- Hubble CLI: queries flows (
hubble observe) filtered by pod, namespace, identity, verdict, FQDN or L7 protocol. 1.20 added--reply/--not-replyfilters.
Hubble sees DNS queries, HTTP requests and responses, gRPC calls, TCP flags and policy drop reasons without application changes or sidecars. L7 detail requires the flow to go through the proxy (an L7 policy or visibility rule). Since 1.18, policies can carry a free-text log field that shows up in Hubble flows, so a flow can be traced back to the rule that allowed or denied it.
Hubble Flow Visibility¶
Hubble matters for security auditing and incident response:
- Flow logging: each flow is recorded with source/destination identity, verdict (forwarded/denied/dropped), L4 protocol, and L7 metadata (HTTP method/path, DNS query, gRPC method).
- Policy drop visibility: denied flows carry the drop reason, which makes policy debugging much faster.
- DNS monitoring: DNS queries and responses are logged, which helps detect data exfiltration and unexpected domains.
- Service map: Hubble UI draws service-to-service traffic and highlights denied flows.
- Hubble metrics: Prometheus metrics for flow rates, drops per namespace, DNS and HTTP stats.
- Hubble exporter: writes flows to files for external SIEM/logging pipelines and long-term retention.
Tetragon¶
Tetragon is a separate eBPF-based runtime security project under the Cilium umbrella (CNCF lists it as a Cilium sub-project). It runs as its own DaemonSet, does not depend on the Cilium CNI datapath, and works on clusters with other CNIs and on plain Linux hosts. Latest Helm chart: 1.7.1 (2026-08-25).
- Process execution monitoring: traces
execve,exitand process ancestry. - File access monitoring: file open, read, write and permission changes through kprobes/LSM hooks.
- Network monitoring: TCP connect, accept, bind and close at the socket level.
- TracingPolicy CRDs (
cilium.io/v1alpha1): declare which kernel functions to hook, with in-kernel filters (matchArgs,matchBinaries,matchNamespacesand others). - In-kernel enforcement: actions such as
SigkillorOverride(change a return value) run synchronously in the kernel, so the process is stopped before the operation finishes rather than after a userspace agent notices.
Tetragon attaches to kprobes, tracepoints and LSM hooks, not to TC/XDP. Filtering in the kernel keeps event volume and overhead low. This diagram shows how a TracingPolicy becomes kernel programs and actions.
flowchart LR
subgraph Kernel_T["Kernel"]
LSM["LSM hooks"]
Kprobes["kprobes / tracepoints / uprobes"]
eBPF_T["Tetragon eBPF programs<br/>(in-kernel filters)"]
end
subgraph Userspace_T["Tetragon agent"]
PolicyEngine["Policy loader<br/>(TracingPolicy CRDs)"]
EventProc["Event processor<br/>(adds pod / process context)"]
end
subgraph Actions["Outcomes"]
Log["JSON events / gRPC export"]
Kill["Sigkill or Override<br/>(in kernel)"]
Metrics["Prometheus metrics"]
end
PolicyEngine --> eBPF_T
LSM --> eBPF_T
Kprobes --> eBPF_T
eBPF_T --> Kill
eBPF_T --> EventProc
EventProc --> Log
EventProc --> Metrics
style Kernel_T fill:#c62828,color:#fff
Tetragon Use Cases¶
| Use Case | TracingPolicy Target |
|---|---|
| Detect container escape | Monitor setns, unshare syscalls |
| Block crypto mining | Kill processes connecting to mining pools |
| File integrity | Alert on writes to /etc/passwd, /etc/shadow |
| Network forensics | Log all TCP connections from a namespace |
| Privilege escalation | Detect setuid(0) calls |
A worked TracingPolicy example is in How-to Guides.
ClusterMesh¶
Cluster Mesh extends Cilium networking across Kubernetes clusters:
- Pod-to-pod connectivity across clusters in a flat IP space (pod CIDRs must not overlap), forwarded directly between nodes without a gateway or proxy.
- Global services and affinity: a Service with the same name and namespace in several clusters, marked global, gets backends from all of them. Affinity can prefer local or remote backends. The Kubernetes Multi-Cluster Services API (MCS-API) is stable since 1.20.
- Cluster-aware policy: identities are shared across clusters, so policies can select remote workloads. 1.20 adds a
cluster-meshentity that selects all endpoints in all meshed clusters.
Each cluster runs clustermesh-apiserver pods with three containers: an embedded etcd that holds the local state exposed to peers over mTLS (no persistent storage, it is rebuilt from kube-apiserver), an apiserver that syncs local Cilium and Kubernetes state into that etcd, and kvstoremesh, which caches remote clusters' state locally so agents do not all connect to every remote cluster. This diagram shows state exchange between two clusters.
graph LR
subgraph A["Cluster A"]
A_KAPI["kube-apiserver"]
A_CMAPI["clustermesh-apiserver<br/>etcd + apiserver + kvstoremesh"]
A_AGENT["cilium-agents"]
end
subgraph B["Cluster B"]
B_KAPI["kube-apiserver"]
B_CMAPI["clustermesh-apiserver<br/>etcd + apiserver + kvstoremesh"]
B_AGENT["cilium-agents"]
end
A_KAPI -->|"local state"| A_CMAPI
B_KAPI -->|"local state"| B_CMAPI
A_CMAPI <-->|"mTLS, remote state sync"| B_CMAPI
A_CMAPI -->|"cached remote state"| A_AGENT
B_CMAPI -->|"cached remote state"| B_AGENT
A_AGENT <-.->|"pod-to-pod traffic, direct"| B_AGENT
One trust domain
All clusters in a mesh form a single trust domain. A compromised Cilium control plane in one cluster can push malicious state to the others. Only connect clusters with equal security posture. Limits (255 or 511 clusters) are in Reference.
kube-proxy Replacement¶
Cilium can replace kube-proxy and handle all Kubernetes service types in eBPF. The per-service-type mapping is in Reference: kube-proxy Replacement Coverage.
Why it is faster
Connect-time load balancing intercepts connect() at the socket level and resolves the ClusterIP to a backend pod IP once. In-cluster traffic then needs no per-packet NAT and no conntrack entries in the kernel's netfilter table, which removes a major source of conntrack pressure at scale.
Modes: kubeProxyReplacement=true gives full replacement (the agent fails to start if the kernel lacks the needed features). false (the Helm default) still load-balances ClusterIP traffic per packet in eBPF but leaves NodePort/LoadBalancer to kube-proxy. The old strict/partial/probe values have been removed. Without kube-proxy the agent needs k8sServiceHost/k8sServicePort (or k8s.apiServerURLs, which since 1.18 supports failover between several API servers).
NodePort modes: SNAT (default), DSR (keeps the client source IP and skips the return hop; dispatch through an IP option/extension header or Geneve), or hybrid (DSR for TCP, SNAT for UDP). Maglev consistent hashing (loadBalancer.algorithm=maglev) applies to north-south traffic and keeps backend selection stable across nodes.
1.20 behavior change
With socket LB disabled or socketLB.hostNamespaceOnly=true, in-cluster connections from pods to NodePort services are now load-balanced when they leave the client pod, not at the target node. Client egress and backend ingress policies must allow that traffic.
Service Mesh and Gateway API¶
Cilium's "sidecar-free" service mesh combines three parts: eBPF for L3/L4 (load balancing, policy, encryption), a per-node Envoy (cilium-envoy) for L7 features, and Kubernetes APIs (Ingress, Gateway API, GAMMA, CiliumEnvoyConfig) as the configuration surface.
- Gateway API: Cilium 1.20 supports Gateway API v1.6.1 (it jumped from v1.4 in 1.19), including TLSRoute v1, TCPRoute, UDPRoute, ListenerSets, BackendTLSPolicy and the ExternalAuth filter (GEP-1494, gRPC or HTTP ext_authz).
- GAMMA (east-west Gateway API): HTTPRoute and, since 1.19, GRPCRoute.
- mTLS: the beta SPIFFE-based "mutual authentication" feature is deprecated in 1.20. The suggested path is ztunnel transparent encryption (beta since 1.19), which uses the Istio ztunnel proxy (image
quay.io/cilium/ztunnelsince 1.20) for per-node L4 mTLS over HBONE (port 15008), with SPIRE or an internal CA.
Traffic through the proxy is disrupted during Cilium upgrades (connections must reconnect). L3/L4-only traffic is not.
Identity-Based Security Model¶
Cilium's security model is built on identities, not IP addresses. Pods with the same security-relevant labels share a numeric identity, and policy is enforced on identities at the receiving side:
- Identity allocation: the agent computes a numeric identity from a pod's label set. Pods with identical labels share it. Identities are stored as
CiliumIdentityCRDs (or in etcd in kvstore mode) and synchronized across nodes and clusters. - Identity on the wire: in tunnel mode the source identity is carried in the VXLAN/Geneve header. In native routing the receiver derives it from the source IP through the IP cache map. Between endpoints on the same node it is passed in packet metadata.
- Policy enforcement: eBPF programs compare the source identity (plus port and protocol) against the destination endpoint's policy map.
- Special identities: reserved identities such as
host,remote-node,world,kube-apiserverandhealthlet policies refer to non-pod traffic through entities.
Advantage over IP-based policies
Identity-based policy keeps working through IP changes from pod restarts, scaling and rescheduling. Rules stay valid as long as labels do not change.
Network Policy CRDs¶
Kubernetes NetworkPolicy¶
Cilium fully implements the Kubernetes NetworkPolicy resource and translates it into eBPF policy. Since 1.20 it also supports the SIG Network Policy API's cluster-scoped ClusterNetworkPolicy (KCNP).
CiliumNetworkPolicy (cilium.io/v2)¶
CiliumNetworkPolicy extends the Kubernetes model with richer matching. This example shows L7 inspection of TLS traffic to an external FQDN together with the DNS rule that toFQDNs needs:
apiVersion: "cilium.io/v2"
kind: CiliumNetworkPolicy
metadata:
name: "l7-visibility-tls"
spec:
description: "L7 policy with TLS visibility"
endpointSelector:
matchLabels:
org: empire
class: mediabot
egress:
- toFQDNs:
- matchName: "httpbin.org"
toPorts:
- ports:
- port: "443"
protocol: "TCP"
terminatingTLS:
secret:
namespace: "kube-system"
name: "httpbin-tls-data"
originatingTLS:
secret:
namespace: "kube-system"
name: "tls-orig-data"
rules:
http:
- {}
- toPorts:
- ports:
- port: "53"
protocol: ANY
rules:
dns:
- matchPattern: "*"
Key capabilities of CiliumNetworkPolicy:
- FQDN-based rules:
toFQDNsallows egress by domain name. The DNS proxy watches DNS answers and maps names to IPs, so a matching L7dnsrule is required. Since 1.19**.matches any depth of subdomains (**.cilium.iomatchesa.b.cilium.iobut notcilium.io). - L7 HTTP/gRPC rules: method, path, header and host matching in Envoy. gRPC is matched as HTTP/2 paths. Kafka L7 rules were removed in 1.20 (deprecated since 1.18).
- DNS-aware rules:
rules.dnsfilters which names pods may resolve. - TLS interception:
originatingTLSandterminatingTLSlet Envoy decrypt and re-encrypt traffic with stored certificates (a man-in-the-middle by design) so L7 rules can apply. - CIDR rules:
toCIDRSet/fromCIDRSetandCiliumCIDRGroup(v2 since 1.18). Since 1.20 CIDR selectors can also match in-cluster IPs. - Entity rules:
toEntities/fromEntitiesforhost,remote-node,cluster,cluster-mesh(1.20),world,kube-apiserverandall. - Deny rules:
ingressDeny/egressDenyalways win over allow rules, whatever policy type the allow comes from. enableDefaultDeny: lets a policy select endpoints without switching them into default-deny, which is useful for cluster-wide DNS interception or audit policies.
CiliumClusterwideNetworkPolicy¶
A cluster-scoped variant that applies across all namespaces, comparable to Calico's GlobalNetworkPolicy. It is the usual place for baseline policies such as default-deny, DNS allow rules, and blocks on world ingress.
Default Deny¶
Cilium allows all traffic to an endpoint until a policy selects it. As soon as any rule with an ingress (or egress) section selects an endpoint, that direction switches to default deny and only explicitly allowed traffic passes. So "default deny" in Cilium means "select everything with a policy that allows little or nothing". It does not mean a deny rule against all. Deny rules take precedence over every allow, so such a policy could never be relaxed. Working manifests are in How-to Guides.
Transparent Encryption¶
Cilium supports three modes of transparent encryption for pod-to-pod (and optionally node-to-node) traffic:
| Mode | How it works | Requirements |
|---|---|---|
| WireGuard | In-kernel WireGuard device (cilium_wg0). Each node generates a key pair. Public keys are shared through CiliumNode resources. |
Kernel 5.6+ (CONFIG_WIREGUARD) or the out-of-tree module. UDP 51871 open between nodes |
| IPsec | Linux XFRM with a pre-shared key from the cilium-ipsec-keys Secret. Rotation is manual: you replace the Secret, and agents apply the new key with random jitter. |
Kernel XFRM support. ESP allowed through firewalls |
| ztunnel (Beta, 1.19+) | Per-node ztunnel proxy does L4 mTLS (HBONE) for workloads in namespaces labeled io.cilium/mtls-enabled=true. Identities come from SPIRE or an internal CA. |
encryption.type=ztunnel. No sidecars |
Both IPsec and WireGuard support a strict mode since 1.19 that drops unencrypted traffic between nodes instead of sending it in clear. Helm commands are in How-to Guides.
WireGuard vs IPsec
WireGuard is the usual choice for new deployments. It has a smaller codebase, simpler key management (no shared secret to rotate), and good performance with ChaCha20-Poly1305. IPsec stays relevant where compliance requires FIPS-validated ciphers (AES-GCM) or where IPsec infrastructure already exists.
Tetragon Runtime Security¶
Tetragon complements Cilium's network security at the process level. Network policy decides which identities may talk. Tetragon watches what processes inside those workloads actually do (exec, file access, privilege changes, sockets) and can kill them in the kernel. See Tetragon for the architecture and How-to Guides for a policy example. Tetragon runs independently of Cilium's CNI datapath and can be deployed in clusters that use other CNI plugins.
Threat Model¶
| Threat | Mitigation |
|---|---|
| Lateral pod movement | Default-deny baseline, identity-based allow rules |
| Unauthorized egress | FQDN-based egress policies, DNS filtering, CIDR rules, egress gateway |
| Unencrypted inter-node traffic | WireGuard or IPsec (with strict mode) or ztunnel mTLS |
| Pod identity spoofing | Identity is derived by Cilium from labels and the IP cache. Pods cannot set it. Source IP verification is on by default (1.20 allows a per-pod opt-out by annotation) |
| Compromised container | Tetragon runtime enforcement (process kill, file access alerts) |
| Data exfiltration | DNS monitoring, L7 policy, Hubble flow logging |
| Supply chain / compromised image | Network isolation + Tetragon process monitoring |
| Blind spots after incidents | Hubble flow logging with policy verdicts and the policy log field |
| Compromised peer cluster in Cluster Mesh | None inside Cilium: the mesh is one trust domain. Only mesh clusters with equal security posture |
| Privileged agent compromise | cilium-agent needs CAP_SYS_ADMIN and the host network. Protect the kube-system namespace and node access |
Sources¶
- Cilium component overview
- Cilium eBPF datapath (code overview)
- eBPF & XDP reference
- kube-proxy replacement
- Hubble observability
- IPAM modes
- Cluster Mesh architecture
- netkit device mode (tuning guide)
- Network policy overview
- Deny policies
- Transparent encryption
- TLS visibility
- Cilium threat model
- Tetragon documentation
- Cilium v1.20 CHANGELOG
- Native mTLS for Cilium with ztunnel (cilium.io blog, 2026-03-23)