Skip to content

Explanation

What this page covers

How Cilium works and why it is built this way: agent, operator and CNI roles, the eBPF datapath (TC, XDP, socket hooks, veth vs netkit), identity-based policy, Hubble, Tetragon, Cluster Mesh, kube-proxy replacement, encryption, and the threat model. Look-up tables (ports, Helm values, map limits, IPAM and routing modes, benchmarks) are in Reference. Step-by-step tasks are in How-to Guides. Topic hub: Cilium.

Overview

Cilium is an eBPF-based CNI plugin for Kubernetes that provides networking, security, and observability. Traditional CNI plugins rely on iptables. Cilium instead runs eBPF programs inside the Linux kernel for packet forwarding, policy enforcement, load balancing, and tracing. It supports tunnel (VXLAN/Geneve) and native routing modes, and it can replace kube-proxy entirely.

The design has three main ideas:

  1. Programmable kernel datapath. The agent compiles and loads eBPF programs and keeps state in BPF maps. Lookups are hash-map based, not linear rule chains.
  2. Identity instead of IP. Pods get a numeric security identity derived from their labels. Policy is enforced on identities, so it survives pod churn and IP reuse.
  3. Visibility from the same hooks. The programs that forward and filter packets also emit flow events, which Hubble turns into flow logs, service maps and metrics.

Component Diagram

This diagram shows the Cilium 1.20 control plane and per-node components and how they relate to Kubernetes, Hubble, Cluster Mesh and Tetragon.

graph TB
    subgraph K8S["Kubernetes control plane"]
        KAPI["kube-apiserver<br/>(Pods, Services, EndpointSlices,<br/>CNP/CCNP, CiliumIdentity CRDs)"]
    end

    subgraph OPS["cilium-operator (Deployment, 2 replicas)"]
        OPER["IPAM (cluster-pool, multi-pool, ENI, Azure),<br/>CRD registration, identity GC,<br/>Gateway API / Ingress controller"]
    end

    subgraph NODE["Each node"]
        AGENT["cilium-agent (DaemonSet)<br/>compiles + loads eBPF,<br/>allocates identities, embeds Hubble server"]
        ENVOY["cilium-envoy (DaemonSet)<br/>L7 policy, Ingress, Gateway API"]
        CNI["cilium-cni binary<br/>(called by kubelet)"]
        BPF["eBPF programs<br/>TC/tcx, XDP, cgroup socket hooks"]
        MAPS["BPF maps<br/>ipcache, policy, lb, CT, NAT"]
    end

    subgraph OBS["Observability"]
        RELAY["Hubble Relay :4245"]
        UI["Hubble UI"]
    end

    subgraph CM["Cluster Mesh (optional)"]
        CMAPI["clustermesh-apiserver<br/>etcd + apiserver + kvstoremesh"]
    end

    TETRA["Tetragon (separate DaemonSet)<br/>kprobes, tracepoints, LSM"]

    KAPI -->|"watch"| AGENT
    KAPI -->|"watch"| OPER
    OPER -->|"pod CIDRs via CiliumNode"| AGENT
    CNI -->|"create endpoint"| AGENT
    AGENT -->|"load"| BPF
    AGENT -->|"write"| MAPS
    BPF -->|"read"| MAPS
    BPF -->|"redirect L7 traffic"| ENVOY
    AGENT -->|"flows :4244"| RELAY
    RELAY --> UI
    CMAPI -->|"remote cluster state"| AGENT
    KAPI -->|"local state"| CMAPI

Core Components

Cilium Agent

The Cilium agent (cilium-agent) runs as a DaemonSet on every node. It is the per-node control plane:

  • eBPF program management: compiles, loads, and attaches eBPF programs to TC (tcx on newer kernels), XDP, and cgroup socket hooks. It regenerates and swaps programs when endpoints or policies change.
  • Endpoint tracking: watches the Kubernetes API for pods, creates Cilium endpoints, and assigns security identities.
  • Identity allocation: pods with the same security-relevant labels share one numeric identity. Identities are stored as CiliumIdentity CRDs (default) or in etcd (kvstore mode).
  • Service load balancing: programs BPF maps for ClusterIP, NodePort, LoadBalancer, and ExternalIP services. In kube-proxy replacement mode this fully replaces iptables service routing. The load-balancing control plane was redesigned in 1.18 to use less memory, and 1.20 reworked the backend representation so thousands of services can share a backend efficiently.
  • Policy enforcement: turns Kubernetes NetworkPolicy, CiliumNetworkPolicy, CiliumClusterwideNetworkPolicy and (since 1.20) Kubernetes ClusterNetworkPolicy into per-endpoint BPF policy maps.
  • Routing: programs kernel routes for native routing, or manages VXLAN/Geneve devices for overlay mode.
  • Embedded services: the Hubble server (port 4244), the DNS proxy used for toFQDNs, and the BGP control plane (GoBGP v4 since 1.20).

Cilium Operator

The Cilium Operator runs as a Deployment (2 replicas by default) and handles cluster-wide jobs:

  • IPAM management: allocates pod CIDRs to nodes in cluster-pool and multi-pool mode, and cloud IPs/prefixes in ENI, Azure and AlibabaCloud modes.
  • CRD lifecycle: registers and maintains CiliumIdentity, CiliumEndpoint/CiliumEndpointSlice, CiliumNode and other CRDs.
  • Garbage collection: removes stale identities and endpoints.
  • Controllers: runs the Gateway API and Ingress controllers that turn Kubernetes resources into Envoy configuration, and LB-IPAM for LoadBalancer IPs.

CNI Plugin

The CNI binary (cilium-cni) is called by the kubelet (through the container runtime) when a pod is created. It:

  1. Calls the Cilium agent API to create an endpoint for the pod.
  2. Creates the pod device (a veth pair by default, or a netkit pair) and moves one end into the pod's network namespace.
  3. Waits for the agent to regenerate the endpoint's eBPF programs and policy so the pod starts with the correct rules.

Since 1.20 Cilium uses CNI spec version 1.0.0 by default.

cilium-envoy

L7 functions (HTTP/gRPC policy, TLS interception, Ingress, Gateway API, GAMMA) are handled by Envoy. For new installs Envoy runs as a separate cilium-envoy DaemonSet (envoy.enabled=true) rather than inside the agent pod, so it can be upgraded and resourced on its own. There is one Envoy per node, not a sidecar per pod. eBPF redirects only the flows that need L7 handling to it. Cilium 1.20 ships cilium-envoy v1.37.x and adds an ADS (Aggregated Discovery Service) server.

clustermesh-apiserver

With Cluster Mesh enabled, each cluster runs clustermesh-apiserver pods that expose local state to peer clusters and pull remote state in. See ClusterMesh.


eBPF Datapath

Cilium's datapath is built on eBPF. Programs are attached at several hook points in the kernel networking stack. The diagram shows where each hook sits on ingress and egress and which maps it consults.

graph TD
    subgraph ING["Packet flow: ingress path"]
        NIC["Network Interface<br/>(eth0)"]
        XDP["XDP Hook<br/>(DDoS mitigation,<br/>early drop, LB acceleration)"]
        TC_ING["TC ingress on host device<br/>(policy check,<br/>load balancing, routing)"]
        L3["eBPF host routing<br/>(bpf_redirect_neigh / redirect_peer)"]
        VETH["Pod device<br/>(veth or netkit)"]
    end

    subgraph EGR["Packet flow: egress path"]
        POD_EG["Pod sends packet"]
        TC_EG["TC hook on pod device<br/>(policy check, NAT, SNAT)"]
        SOCK_LB["Socket-level LB<br/>cgroup connect/sendmsg hooks<br/>(connect-time LB, no per-packet NAT)"]
    end

    subgraph MAPS["BPF maps (shared state)"]
        EP_MAP["Endpoint / ipcache maps<br/>IP -> identity"]
        SVC_MAP["Service maps<br/>frontend -> backends"]
        POL_MAP["Policy map<br/>identity + port -> verdict"]
        CT_MAP["Connection tracking map"]
    end

    NIC -->|"raw packet"| XDP
    XDP -->|"pass"| TC_ING
    TC_ING -->|"lookup"| EP_MAP
    TC_ING -->|"service resolve"| SVC_MAP
    TC_ING -->|"policy check"| POL_MAP
    TC_ING -->|"allowed"| L3
    L3 -->|"local delivery"| VETH

    POD_EG -->|"packet"| TC_EG
    POD_EG -->|"connect()"| SOCK_LB
    SOCK_LB -->|"direct to backend"| DEST["Destination Pod"]
    TC_EG -->|"lookup"| CT_MAP
    TC_EG -->|"policy check"| POL_MAP
    TC_EG -->|"forward"| NIC

eBPF Hook Points

Hook Point Attachment Purpose
XDP (eXpress Data Path) Network driver, before skb allocation Early drop, DDoS mitigation, XDP-accelerated NodePort/LoadBalancer (including DSR)
TC / tcx ingress Physical, veth and netkit devices Policy enforcement, service load balancing, routing decisions
TC / tcx egress Physical and pod devices SNAT, policy enforcement, tunnel encapsulation
cgroup socket hooks cgroup v2 root Connect-time load balancing (connect(), sendmsg()), which skips per-packet NAT
netkit peer programs Inside the netkit device (1.20: bpf.datapathMode=netkit) Pod programs run in the device itself, removing the veth hop

When netkit is enabled, Cilium uses tcx (BPF links) for attachments on all other devices too.

BPF Maps

BPF maps are the shared data structures between the Cilium agent (userspace) and the eBPF programs (kernel):

  • Endpoint map: local pod IPs to endpoint ID, MAC and interface index.
  • IP cache: every known IP (local and remote, across clusters) to its security identity and tunnel endpoint.
  • Service maps: ClusterIP/NodePort/LoadBalancer frontends and their backends.
  • Policy map (per endpoint): allowed identity + port + protocol entries. 1.20 added wildcard entries so the world, remote-node, cluster and new cluster-mesh entities use far fewer entries.
  • Connection tracking and NAT maps: stateful policy and NAT state.
  • Tunnel map: remote node IPs to tunnel endpoints (VXLAN/Geneve mode).

BPF map memory

BPF maps are created with fixed upper limits and allocated in kernel memory. At large scale the CT, NAT and IP cache maps are the big consumers. Size them with bpf.mapDynamicSizeRatio (default 0.25% of node memory) or explicit values such as bpf.ctTcpMax, bpf.ctAnyMax, bpf.natMax and bpf.policyMapMax. Default limits are listed in Reference: eBPF Map Limits.

veth vs netkit

By default each pod gets a veth pair. A packet leaving the pod crosses the veth into the host namespace, where Cilium's TC program handles it. netkit (Linux 6.8+) is a device type built for Cilium: the pod's BPF program runs inside the netkit peer itself. With eBPF host routing this makes the namespace switch almost free, so pods get near host-namespace throughput and latency. It also supports BIG TCP.

Status in 1.20: beta. bpf.datapathMode accepts veth (default), netkit (L3), netkit-l2, and auto (new in 1.20: use netkit if the kernel supports it, otherwise veth). netkit requires eBPF host routing and does not work with bpf.tproxy. Existing pods cannot switch in place: veth and netkit cannot be mixed on one node. Enable it on new nodes (per-node config) or drain nodes first. The Cilium docs say veth mode will be deprecated once 6.8+ kernels are common.


How It Works

eBPF data plane internals, packet flow, the Hubble observability pipeline, and Tetragon enforcement.

eBPF Data Plane

Traditional CNIs use iptables, a linear chain of rules that gets slower as rules grow (O(n)). Cilium replaces this with eBPF hash maps that give constant-time O(1) lookups in kernel space. This diagram contrasts the two lookup models.

flowchart LR
    subgraph Traditional["iptables-based CNI"]
        PKT1["Packet"] --> R1["Rule 1"] --> R2["Rule 2"] --> R3["Rule 3"] --> RN["Rule N<br/>(O(n) traversal)"]
    end

    subgraph Cilium_DP["Cilium eBPF Data Plane"]
        PKT2["Packet"] --> MAP["eBPF Hash Map<br/>(O(1) lookup)"] --> Action["Allow / Drop / Redirect"]
    end

    style Traditional fill:#c62828,color:#fff
    style Cilium_DP fill:#2e7d32,color:#fff

Packet Flow — Pod-to-Pod (Same Node)

With eBPF host routing, a packet between two pods on the same node never goes through the host's iptables or routing stack. The sequence shows the veth case.

sequenceDiagram
    participant PodA as Pod A
    participant LXC_A as lxc device of Pod A (host side)
    participant BPF as Cilium eBPF (from-container)
    participant Maps as ipcache + policy maps
    participant PodB as Pod B

    PodA->>LXC_A: Send packet
    LXC_A->>BPF: TC hook runs from-container program
    BPF->>Maps: Resolve destination identity, check egress policy of A
    BPF->>Maps: Check ingress policy of B, create CT entry
    BPF->>PodB: bpf_redirect_peer into Pod B namespace
    Note over BPF,PodB: Host network stack and iptables are bypassed

Packet Flow — Pod-to-Service (ClusterIP)

With socket-level load balancing (enabled by kube-proxy replacement), the service is resolved once at connect(). Packets then carry the backend IP from the start, so no per-packet DNAT is needed.

sequenceDiagram
    participant Pod as Client Pod
    participant Sock as cgroup connect hook
    participant SVC as Service map
    participant TC as TC eBPF (from-container)
    participant Backend as Backend Pod

    Pod->>Sock: connect() to ClusterIP:port
    Sock->>SVC: Look up frontend, pick backend (random or Maglev)
    Sock-->>Pod: Socket now bound to backend IP:port
    Pod->>TC: Packets addressed to backend
    TC->>Backend: Policy check, then forward directly
    Note over Sock,SVC: Resolved once per connection, not per packet

Hubble Observability Pipeline

Datapath programs emit events to the agent through a perf event buffer (ring buffer support was added in 1.18). The Hubble server inside each agent turns them into flows. Relay aggregates flows cluster-wide. Metrics and flow export come from each agent.

flowchart TB
    subgraph Kernel["Kernel space"]
        eBPF_H["eBPF programs<br/>(TC, socket, XDP)"]
        PerfBuf["Perf event buffer<br/>(trace, drop, policy verdict events)"]
    end

    subgraph Userspace["Hubble stack"]
        HubbleAgent["Hubble server :4244<br/>(embedded in cilium-agent)"]
        HubbleRelay["Hubble Relay :4245<br/>(cluster-wide aggregation)"]
        HubbleUI["Hubble UI<br/>(service map, flow table)"]
        CLI["hubble CLI"]
    end

    subgraph Export["Export"]
        Prom["Prometheus<br/>(Hubble metrics per agent)"]
        SIEM["Flow exporter<br/>(JSON file for log / SIEM pipelines)"]
    end

    eBPF_H -->|"perf events"| PerfBuf
    PerfBuf --> HubbleAgent
    HubbleAgent -->|"gRPC"| HubbleRelay
    HubbleRelay --> HubbleUI
    HubbleRelay --> CLI
    HubbleAgent --> Prom
    HubbleAgent --> SIEM

    style Kernel fill:#f9a825,color:#000
    style Userspace fill:#7b1fa2,color:#fff

Hubble

Hubble is Cilium's built-in network observability layer:

  • Hubble server: embedded in the Cilium agent. Serves flow events (L3/L4/L7 metadata, policy verdicts) over gRPC on port 4244.
  • Hubble Relay: a Deployment that connects to every node's Hubble server and serves a cluster-wide gRPC API on port 4245.
  • Hubble UI: web interface with a service dependency map, flow table, and policy view.
  • Hubble CLI: queries flows (hubble observe) filtered by pod, namespace, identity, verdict, FQDN or L7 protocol. 1.20 added --reply/--not-reply filters.

Hubble sees DNS queries, HTTP requests and responses, gRPC calls, TCP flags and policy drop reasons without application changes or sidecars. L7 detail requires the flow to go through the proxy (an L7 policy or visibility rule). Since 1.18, policies can carry a free-text log field that shows up in Hubble flows, so a flow can be traced back to the rule that allowed or denied it.

Hubble Flow Visibility

Hubble matters for security auditing and incident response:

  • Flow logging: each flow is recorded with source/destination identity, verdict (forwarded/denied/dropped), L4 protocol, and L7 metadata (HTTP method/path, DNS query, gRPC method).
  • Policy drop visibility: denied flows carry the drop reason, which makes policy debugging much faster.
  • DNS monitoring: DNS queries and responses are logged, which helps detect data exfiltration and unexpected domains.
  • Service map: Hubble UI draws service-to-service traffic and highlights denied flows.
  • Hubble metrics: Prometheus metrics for flow rates, drops per namespace, DNS and HTTP stats.
  • Hubble exporter: writes flows to files for external SIEM/logging pipelines and long-term retention.

Tetragon

Tetragon is a separate eBPF-based runtime security project under the Cilium umbrella (CNCF lists it as a Cilium sub-project). It runs as its own DaemonSet, does not depend on the Cilium CNI datapath, and works on clusters with other CNIs and on plain Linux hosts. Latest Helm chart: 1.7.1 (2026-08-25).

  • Process execution monitoring: traces execve, exit and process ancestry.
  • File access monitoring: file open, read, write and permission changes through kprobes/LSM hooks.
  • Network monitoring: TCP connect, accept, bind and close at the socket level.
  • TracingPolicy CRDs (cilium.io/v1alpha1): declare which kernel functions to hook, with in-kernel filters (matchArgs, matchBinaries, matchNamespaces and others).
  • In-kernel enforcement: actions such as Sigkill or Override (change a return value) run synchronously in the kernel, so the process is stopped before the operation finishes rather than after a userspace agent notices.

Tetragon attaches to kprobes, tracepoints and LSM hooks, not to TC/XDP. Filtering in the kernel keeps event volume and overhead low. This diagram shows how a TracingPolicy becomes kernel programs and actions.

flowchart LR
    subgraph Kernel_T["Kernel"]
        LSM["LSM hooks"]
        Kprobes["kprobes / tracepoints / uprobes"]
        eBPF_T["Tetragon eBPF programs<br/>(in-kernel filters)"]
    end

    subgraph Userspace_T["Tetragon agent"]
        PolicyEngine["Policy loader<br/>(TracingPolicy CRDs)"]
        EventProc["Event processor<br/>(adds pod / process context)"]
    end

    subgraph Actions["Outcomes"]
        Log["JSON events / gRPC export"]
        Kill["Sigkill or Override<br/>(in kernel)"]
        Metrics["Prometheus metrics"]
    end

    PolicyEngine --> eBPF_T
    LSM --> eBPF_T
    Kprobes --> eBPF_T
    eBPF_T --> Kill
    eBPF_T --> EventProc
    EventProc --> Log
    EventProc --> Metrics

    style Kernel_T fill:#c62828,color:#fff

Tetragon Use Cases

Use Case TracingPolicy Target
Detect container escape Monitor setns, unshare syscalls
Block crypto mining Kill processes connecting to mining pools
File integrity Alert on writes to /etc/passwd, /etc/shadow
Network forensics Log all TCP connections from a namespace
Privilege escalation Detect setuid(0) calls

A worked TracingPolicy example is in How-to Guides.


ClusterMesh

Cluster Mesh extends Cilium networking across Kubernetes clusters:

  • Pod-to-pod connectivity across clusters in a flat IP space (pod CIDRs must not overlap), forwarded directly between nodes without a gateway or proxy.
  • Global services and affinity: a Service with the same name and namespace in several clusters, marked global, gets backends from all of them. Affinity can prefer local or remote backends. The Kubernetes Multi-Cluster Services API (MCS-API) is stable since 1.20.
  • Cluster-aware policy: identities are shared across clusters, so policies can select remote workloads. 1.20 adds a cluster-mesh entity that selects all endpoints in all meshed clusters.

Each cluster runs clustermesh-apiserver pods with three containers: an embedded etcd that holds the local state exposed to peers over mTLS (no persistent storage, it is rebuilt from kube-apiserver), an apiserver that syncs local Cilium and Kubernetes state into that etcd, and kvstoremesh, which caches remote clusters' state locally so agents do not all connect to every remote cluster. This diagram shows state exchange between two clusters.

graph LR
    subgraph A["Cluster A"]
        A_KAPI["kube-apiserver"]
        A_CMAPI["clustermesh-apiserver<br/>etcd + apiserver + kvstoremesh"]
        A_AGENT["cilium-agents"]
    end

    subgraph B["Cluster B"]
        B_KAPI["kube-apiserver"]
        B_CMAPI["clustermesh-apiserver<br/>etcd + apiserver + kvstoremesh"]
        B_AGENT["cilium-agents"]
    end

    A_KAPI -->|"local state"| A_CMAPI
    B_KAPI -->|"local state"| B_CMAPI
    A_CMAPI <-->|"mTLS, remote state sync"| B_CMAPI
    A_CMAPI -->|"cached remote state"| A_AGENT
    B_CMAPI -->|"cached remote state"| B_AGENT
    A_AGENT <-.->|"pod-to-pod traffic, direct"| B_AGENT

One trust domain

All clusters in a mesh form a single trust domain. A compromised Cilium control plane in one cluster can push malicious state to the others. Only connect clusters with equal security posture. Limits (255 or 511 clusters) are in Reference.


kube-proxy Replacement

Cilium can replace kube-proxy and handle all Kubernetes service types in eBPF. The per-service-type mapping is in Reference: kube-proxy Replacement Coverage.

Why it is faster

Connect-time load balancing intercepts connect() at the socket level and resolves the ClusterIP to a backend pod IP once. In-cluster traffic then needs no per-packet NAT and no conntrack entries in the kernel's netfilter table, which removes a major source of conntrack pressure at scale.

Modes: kubeProxyReplacement=true gives full replacement (the agent fails to start if the kernel lacks the needed features). false (the Helm default) still load-balances ClusterIP traffic per packet in eBPF but leaves NodePort/LoadBalancer to kube-proxy. The old strict/partial/probe values have been removed. Without kube-proxy the agent needs k8sServiceHost/k8sServicePort (or k8s.apiServerURLs, which since 1.18 supports failover between several API servers).

NodePort modes: SNAT (default), DSR (keeps the client source IP and skips the return hop; dispatch through an IP option/extension header or Geneve), or hybrid (DSR for TCP, SNAT for UDP). Maglev consistent hashing (loadBalancer.algorithm=maglev) applies to north-south traffic and keeps backend selection stable across nodes.

1.20 behavior change

With socket LB disabled or socketLB.hostNamespaceOnly=true, in-cluster connections from pods to NodePort services are now load-balanced when they leave the client pod, not at the target node. Client egress and backend ingress policies must allow that traffic.


Service Mesh and Gateway API

Cilium's "sidecar-free" service mesh combines three parts: eBPF for L3/L4 (load balancing, policy, encryption), a per-node Envoy (cilium-envoy) for L7 features, and Kubernetes APIs (Ingress, Gateway API, GAMMA, CiliumEnvoyConfig) as the configuration surface.

  • Gateway API: Cilium 1.20 supports Gateway API v1.6.1 (it jumped from v1.4 in 1.19), including TLSRoute v1, TCPRoute, UDPRoute, ListenerSets, BackendTLSPolicy and the ExternalAuth filter (GEP-1494, gRPC or HTTP ext_authz).
  • GAMMA (east-west Gateway API): HTTPRoute and, since 1.19, GRPCRoute.
  • mTLS: the beta SPIFFE-based "mutual authentication" feature is deprecated in 1.20. The suggested path is ztunnel transparent encryption (beta since 1.19), which uses the Istio ztunnel proxy (image quay.io/cilium/ztunnel since 1.20) for per-node L4 mTLS over HBONE (port 15008), with SPIRE or an internal CA.

Traffic through the proxy is disrupted during Cilium upgrades (connections must reconnect). L3/L4-only traffic is not.


Identity-Based Security Model

Cilium's security model is built on identities, not IP addresses. Pods with the same security-relevant labels share a numeric identity, and policy is enforced on identities at the receiving side:

  1. Identity allocation: the agent computes a numeric identity from a pod's label set. Pods with identical labels share it. Identities are stored as CiliumIdentity CRDs (or in etcd in kvstore mode) and synchronized across nodes and clusters.
  2. Identity on the wire: in tunnel mode the source identity is carried in the VXLAN/Geneve header. In native routing the receiver derives it from the source IP through the IP cache map. Between endpoints on the same node it is passed in packet metadata.
  3. Policy enforcement: eBPF programs compare the source identity (plus port and protocol) against the destination endpoint's policy map.
  4. Special identities: reserved identities such as host, remote-node, world, kube-apiserver and health let policies refer to non-pod traffic through entities.

Advantage over IP-based policies

Identity-based policy keeps working through IP changes from pod restarts, scaling and rescheduling. Rules stay valid as long as labels do not change.


Network Policy CRDs

Kubernetes NetworkPolicy

Cilium fully implements the Kubernetes NetworkPolicy resource and translates it into eBPF policy. Since 1.20 it also supports the SIG Network Policy API's cluster-scoped ClusterNetworkPolicy (KCNP).

CiliumNetworkPolicy (cilium.io/v2)

CiliumNetworkPolicy extends the Kubernetes model with richer matching. This example shows L7 inspection of TLS traffic to an external FQDN together with the DNS rule that toFQDNs needs:

apiVersion: "cilium.io/v2"
kind: CiliumNetworkPolicy
metadata:
  name: "l7-visibility-tls"
spec:
  description: "L7 policy with TLS visibility"
  endpointSelector:
    matchLabels:
      org: empire
      class: mediabot
  egress:
  - toFQDNs:
    - matchName: "httpbin.org"
    toPorts:
    - ports:
      - port: "443"
        protocol: "TCP"
      terminatingTLS:
        secret:
          namespace: "kube-system"
          name: "httpbin-tls-data"
      originatingTLS:
        secret:
          namespace: "kube-system"
          name: "tls-orig-data"
      rules:
        http:
        - {}
  - toPorts:
    - ports:
      - port: "53"
        protocol: ANY
      rules:
        dns:
        - matchPattern: "*"

Key capabilities of CiliumNetworkPolicy:

  • FQDN-based rules: toFQDNs allows egress by domain name. The DNS proxy watches DNS answers and maps names to IPs, so a matching L7 dns rule is required. Since 1.19 **. matches any depth of subdomains (**.cilium.io matches a.b.cilium.io but not cilium.io).
  • L7 HTTP/gRPC rules: method, path, header and host matching in Envoy. gRPC is matched as HTTP/2 paths. Kafka L7 rules were removed in 1.20 (deprecated since 1.18).
  • DNS-aware rules: rules.dns filters which names pods may resolve.
  • TLS interception: originatingTLS and terminatingTLS let Envoy decrypt and re-encrypt traffic with stored certificates (a man-in-the-middle by design) so L7 rules can apply.
  • CIDR rules: toCIDRSet / fromCIDRSet and CiliumCIDRGroup (v2 since 1.18). Since 1.20 CIDR selectors can also match in-cluster IPs.
  • Entity rules: toEntities / fromEntities for host, remote-node, cluster, cluster-mesh (1.20), world, kube-apiserver and all.
  • Deny rules: ingressDeny / egressDeny always win over allow rules, whatever policy type the allow comes from.
  • enableDefaultDeny: lets a policy select endpoints without switching them into default-deny, which is useful for cluster-wide DNS interception or audit policies.

CiliumClusterwideNetworkPolicy

A cluster-scoped variant that applies across all namespaces, comparable to Calico's GlobalNetworkPolicy. It is the usual place for baseline policies such as default-deny, DNS allow rules, and blocks on world ingress.

Default Deny

Cilium allows all traffic to an endpoint until a policy selects it. As soon as any rule with an ingress (or egress) section selects an endpoint, that direction switches to default deny and only explicitly allowed traffic passes. So "default deny" in Cilium means "select everything with a policy that allows little or nothing". It does not mean a deny rule against all. Deny rules take precedence over every allow, so such a policy could never be relaxed. Working manifests are in How-to Guides.


Transparent Encryption

Cilium supports three modes of transparent encryption for pod-to-pod (and optionally node-to-node) traffic:

Mode How it works Requirements
WireGuard In-kernel WireGuard device (cilium_wg0). Each node generates a key pair. Public keys are shared through CiliumNode resources. Kernel 5.6+ (CONFIG_WIREGUARD) or the out-of-tree module. UDP 51871 open between nodes
IPsec Linux XFRM with a pre-shared key from the cilium-ipsec-keys Secret. Rotation is manual: you replace the Secret, and agents apply the new key with random jitter. Kernel XFRM support. ESP allowed through firewalls
ztunnel (Beta, 1.19+) Per-node ztunnel proxy does L4 mTLS (HBONE) for workloads in namespaces labeled io.cilium/mtls-enabled=true. Identities come from SPIRE or an internal CA. encryption.type=ztunnel. No sidecars

Both IPsec and WireGuard support a strict mode since 1.19 that drops unencrypted traffic between nodes instead of sending it in clear. Helm commands are in How-to Guides.

WireGuard vs IPsec

WireGuard is the usual choice for new deployments. It has a smaller codebase, simpler key management (no shared secret to rotate), and good performance with ChaCha20-Poly1305. IPsec stays relevant where compliance requires FIPS-validated ciphers (AES-GCM) or where IPsec infrastructure already exists.


Tetragon Runtime Security

Tetragon complements Cilium's network security at the process level. Network policy decides which identities may talk. Tetragon watches what processes inside those workloads actually do (exec, file access, privilege changes, sockets) and can kill them in the kernel. See Tetragon for the architecture and How-to Guides for a policy example. Tetragon runs independently of Cilium's CNI datapath and can be deployed in clusters that use other CNI plugins.


Threat Model

Threat Mitigation
Lateral pod movement Default-deny baseline, identity-based allow rules
Unauthorized egress FQDN-based egress policies, DNS filtering, CIDR rules, egress gateway
Unencrypted inter-node traffic WireGuard or IPsec (with strict mode) or ztunnel mTLS
Pod identity spoofing Identity is derived by Cilium from labels and the IP cache. Pods cannot set it. Source IP verification is on by default (1.20 allows a per-pod opt-out by annotation)
Compromised container Tetragon runtime enforcement (process kill, file access alerts)
Data exfiltration DNS monitoring, L7 policy, Hubble flow logging
Supply chain / compromised image Network isolation + Tetragon process monitoring
Blind spots after incidents Hubble flow logging with policy verdicts and the policy log field
Compromised peer cluster in Cluster Mesh None inside Cilium: the mesh is one trust domain. Only mesh clusters with equal security posture
Privileged agent compromise cilium-agent needs CAP_SYS_ADMIN and the host network. Protect the kube-system namespace and node access

Sources