How-to Guides¶
Scope
Task recipes for Cilium 1.20: install and upgrade, kube-proxy replacement and migration, network policy, Hubble, Gateway API, Cluster Mesh, encryption, netkit, Tetragon, and troubleshooting. Versions, Helm defaults and ports are in Reference. Internals are in Explanation.
Pin versions
Commands below pin 1.20.2, the latest stable release on 2026-09-25. Check https://raw.githubusercontent.com/cilium/cilium/main/stable.txt for newer patches. Only upgrade one minor version at a time.
Deployment Patterns¶
Installation Methods¶
| Method | Use Case | Notes |
|---|---|---|
Helm (https://helm.cilium.io/) |
Production | Most flexible. Recommended for GitOps |
Helm OCI (oci://quay.io/cilium/charts/cilium) |
Production, air-gapped mirrors | Since 1.19. Charts signed with cosign |
cilium CLI (cilium install) |
Dev, quick setup | Wraps the same Helm chart and auto-detects the environment |
| Managed (GKE Dataplane V2, AKS "Azure CNI Powered by Cilium") | Managed Kubernetes | The provider manages the Cilium version and feature set |
| Standalone Tetragon | Runtime security without Cilium CNI | Separate Helm chart cilium/tetragon |
Production Helm Install (kube-proxy-free)¶
helm repo add cilium https://helm.cilium.io/
helm repo update
# API server address is required when kube-proxy is not present
API_SERVER_IP=<control-plane-ip-or-lb>
API_SERVER_PORT=6443
helm install cilium cilium/cilium --version 1.20.2 \
--namespace kube-system \
--set kubeProxyReplacement=true \
--set k8sServiceHost=${API_SERVER_IP} \
--set k8sServicePort=${API_SERVER_PORT} \
--set hubble.relay.enabled=true \
--set hubble.ui.enabled=true \
--set bpf.masquerade=true \
--set ipam.mode=kubernetes
ipam.mode=kubernetes only makes sense when kube-controller-manager runs with --allocate-node-cidrs. Otherwise leave the Helm default (cluster-pool) and set ipam.operator.clusterPoolIPv4PodCIDRList.
The same chart from the OCI registry, with signature verification:
cosign verify \
--certificate-identity-regexp='https://github.com/cilium/cilium/.*' \
--certificate-oidc-issuer=https://token.actions.githubusercontent.com \
quay.io/cilium/charts/cilium:1.20.2
helm install cilium oci://quay.io/cilium/charts/cilium --version 1.20.2 \
--namespace kube-system --set kubeProxyReplacement=true \
--set k8sServiceHost=${API_SERVER_IP} --set k8sServicePort=${API_SERVER_PORT}
The cosign verify identity and issuer values are the ones given in the Cilium Helm install docs ("OCI Registry" section, v1.20).
kube-proxy Replacement¶
Cilium can replace kube-proxy and handle service load balancing in eBPF. Check the result on a node:
kubectl -n kube-system exec ds/cilium -- cilium-dbg status | grep KubeProxyReplacement
# KubeProxyReplacement: True [eth0 (Direct Routing)]
kubectl -n kube-system exec ds/cilium -- cilium-dbg status --verbose | grep -A20 "KubeProxyReplacement Details"
kubectl -n kube-system exec ds/cilium -- cilium-dbg service list
On a new kubeadm cluster, skip kube-proxy at init with kubeadm init --skip-phases=addon/kube-proxy.
Migrate an Existing Cluster from kube-proxy¶
Not an official one-click procedure
The Cilium docs describe per-node configuration (CiliumNodeConfig) and a per-node CNI migration flow, but no step-by-step kube-proxy-to-KPR migration. The pattern below combines those documented pieces. Rehearse it on a staging cluster. The alternative is a maintenance window: set kubeProxyReplacement=true, delete kube-proxy, and restart Cilium on all nodes.
- Run Cilium with
kubeProxyReplacement=falsenext to kube-proxy, and setk8sServiceHost/k8sServicePort. -
Create a
CiliumNodeConfig(cilium.io/v2) that setskube-proxy-replacement: "true"for nodes carrying a migration label: -
Add a node affinity to the kube-proxy DaemonSet so it does not run on nodes with that label.
- For each node:
kubectl cordon,kubectl label node $NODE io.cilium.migration/kube-proxy-replacement=true, delete the Cilium pod on that node so it restarts with the new config, flush kube-proxy's iptables rules on the node (or reboot it), check withcilium-dbg status, thenkubectl uncordon. - When all nodes are done, set
kubeProxyReplacement=truein Helm values, delete the kube-proxy DaemonSet and theCiliumNodeConfig, and remove the labels.
Upgrade Procedure¶
# 1. Read the upgrade notes for the target minor (1.20: Kafka rules removed,
# TLSRoute v1, CiliumNodeConfig v2alpha1 removed)
# 2. Move to the latest patch of the current minor first, then one minor up
# Pre-flight: pre-pull images so the upgrade does not stall on pulls
helm install cilium-preflight cilium/cilium --version 1.20.2 \
--namespace kube-system --set preflight.enabled=true \
--set agent=false --set operator.enabled=false
kubectl -n kube-system get daemonset cilium-pre-flight-check
helm delete cilium-preflight --namespace kube-system
# Upgrade, keeping your values under version control
helm upgrade cilium cilium/cilium --version 1.20.2 \
--namespace kube-system -f cilium-values.yaml
# Post-upgrade validation
cilium status --wait
cilium connectivity test
Avoid --reuse-values across minors
--reuse-values carries old values forward and skips new chart defaults, which often causes trouble across minor versions. Keep an explicit values file instead. L7/proxied connections (L7 policy, Ingress, Gateway API) are reset during the upgrade. L3/L4 traffic is not.
Commands & Recipes¶
Install the CLIs¶
# cilium CLI
CILIUM_CLI_VERSION=$(curl -s https://raw.githubusercontent.com/cilium/cilium-cli/main/stable.txt)
CLI_ARCH=amd64
curl -L --fail --remote-name-all \
https://github.com/cilium/cilium-cli/releases/download/${CILIUM_CLI_VERSION}/cilium-linux-${CLI_ARCH}.tar.gz{,.sha256sum}
sha256sum --check cilium-linux-${CLI_ARCH}.tar.gz.sha256sum
sudo tar xzvfC cilium-linux-${CLI_ARCH}.tar.gz /usr/local/bin
# Hubble CLI
HUBBLE_VERSION=$(curl -s https://raw.githubusercontent.com/cilium/hubble/main/stable.txt)
curl -L --fail --remote-name-all \
https://github.com/cilium/hubble/releases/download/$HUBBLE_VERSION/hubble-linux-amd64.tar.gz{,.sha256sum}
sha256sum --check hubble-linux-amd64.tar.gz.sha256sum
sudo tar xzvfC hubble-linux-amd64.tar.gz /usr/local/bin
Quick Install and Validate¶
# Quick install with the CLI (auto-detects cluster type)
cilium install --version 1.20.2
cilium status --wait
cilium connectivity test
# Or with Helm and Hubble enabled
helm install cilium cilium/cilium --version 1.20.2 \
--namespace kube-system \
--set hubble.enabled=true \
--set hubble.relay.enabled=true \
--set hubble.ui.enabled=true
Hubble Observability¶
# Enable Relay and UI on an existing install
cilium hubble enable --ui
# Port-forward Hubble Relay (localhost:4245)
cilium hubble port-forward &
# Observe flows
hubble observe --namespace default --follow
hubble observe --verdict DROPPED --follow
hubble observe --pod default/myapp --protocol http
hubble observe --to-fqdn "*.amazonaws.com"
hubble observe --from-pod default/frontend --to-pod default/backend --verdict DROPPED
# Service map in the UI
cilium hubble ui # or: kubectl -n kube-system port-forward svc/hubble-ui 12000:80
Network Policies¶
L7 HTTP policy that only allows GET /api/* from frontend to backend:
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
name: allow-api-get
spec:
endpointSelector:
matchLabels:
app: backend
ingress:
- fromEndpoints:
- matchLabels:
app: frontend
toPorts:
- ports:
- port: "8080"
protocol: TCP
rules:
http:
- method: GET
path: "/api/.*"
DNS-based egress policy. toFQDNs only works if DNS traffic also goes through the DNS proxy, which is what the second rule does:
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
name: allow-googleapis
spec:
endpointSelector:
matchLabels:
app: myapp
egress:
- toEndpoints:
- matchLabels:
io.kubernetes.pod.namespace: kube-system
k8s-app: kube-dns
toPorts:
- ports:
- port: "53"
protocol: ANY
rules:
dns:
- matchPattern: "*"
- toFQDNs:
- matchPattern: "**.googleapis.com" # multi-level wildcard, 1.19+
toPorts:
- ports:
- port: "443"
protocol: TCP
Kafka rules removed in 1.20
Policies with rules.kafka, rules.l7 or rules.l7proto must be cleaned up before upgrading to 1.20. Find them with:
kubectl get cnp,ccnp -A -o json | jq -r '.items[] | select(.. | objects | has("kafka") or has("l7proto")) | "\(.metadata.namespace)/\(.metadata.name)"'
Default Deny Baseline¶
Selecting endpoints with a rule that has an empty ingress/egress section switches them to default deny. This namespace-scoped example denies all egress for role: restricted pods (from the Cilium examples):
apiVersion: "cilium.io/v2"
kind: CiliumNetworkPolicy
metadata:
name: "deny-all-egress"
spec:
endpointSelector:
matchLabels:
role: restricted
egress:
- {}
A cluster-wide baseline usually puts all workload pods (outside kube-system) into default deny and still allows DNS, so later namespace policies only add allows:
apiVersion: cilium.io/v2
kind: CiliumClusterwideNetworkPolicy
metadata:
name: baseline-default-deny
spec:
endpointSelector:
matchExpressions:
- key: io.kubernetes.pod.namespace
operator: NotIn
values: ["kube-system"]
ingress:
- {}
egress:
- toEndpoints:
- matchLabels:
io.kubernetes.pod.namespace: kube-system
k8s-app: kube-dns
toPorts:
- ports:
- port: "53"
protocol: ANY
rules:
dns:
- matchPattern: "*"
Do not use deny-all against all
A policy with ingressDeny: [{fromEntities: ["all"]}] and egressDeny: [{toEntities: ["all"]}] blocks everything, and no allow rule can override it, because deny always wins. Use deny rules only for targeted blocks (for example ingress from world).
Gateway API¶
# 1. Install Gateway API v1.6.1 CRDs (standard channel); add tcproutes/udproutes/listenersets if needed
GW=https://raw.githubusercontent.com/kubernetes-sigs/gateway-api/v1.6.1/config/crd/standard
for crd in gatewayclasses gateways httproutes referencegrants grpcroutes backendtlspolicies tlsroutes; do
kubectl apply --server-side -f ${GW}/gateway.networking.k8s.io_${crd}.yaml
done
# 2. Enable the controller (requires kube-proxy replacement and l7Proxy)
helm upgrade cilium cilium/cilium --version 1.20.2 --namespace kube-system \
-f cilium-values.yaml --set gatewayAPI.enabled=true
kubectl -n kube-system rollout restart deployment/cilium-operator
kubectl -n kube-system rollout restart ds/cilium
kubectl get gatewayclass cilium
Existing TLSRoutes
If you already use TLSRoute v1alpha2, back it up and install the experimental v1.6.1 TLSRoute CRD instead of the standard one. See Reference: Gateway API Support.
Cluster Mesh¶
# Each cluster needs a unique name and ID (and non-overlapping pod CIDRs)
cilium install --set cluster.name=$CLUSTER1 --set cluster.id=1 --context $CLUSTER1
cilium install --set cluster.name=$CLUSTER2 --set cluster.id=2 --context $CLUSTER2
cilium clustermesh enable --context $CLUSTER1
cilium clustermesh enable --context $CLUSTER2
cilium clustermesh status --context $CLUSTER1 --wait
cilium clustermesh connect --context $CLUSTER1 --destination-context $CLUSTER2
cilium clustermesh status --context $CLUSTER1 --wait
cilium connectivity test --context $CLUSTER1 --multi-cluster $CLUSTER2
Mark a Service global by annotating it service.cilium.io/global: "true" in each cluster. With Helm-generated certificates, Cluster Mesh certificates expire after one year by default (1.20). Re-render the chart (upgrade) at least yearly, or use the cronJob/certmanager generation methods.
Enable Transparent Encryption¶
# WireGuard (UDP 51871 must be open between nodes)
helm upgrade cilium cilium/cilium --version 1.20.2 --namespace kube-system \
-f cilium-values.yaml \
--set encryption.enabled=true \
--set encryption.type=wireguard
kubectl -n kube-system exec ds/cilium -- cilium-dbg encrypt status
# IPsec: create the pre-shared key first (or: cilium encrypt create-key --auth-algo rfc4106-gcm-aes)
kubectl create -n kube-system secret generic cilium-ipsec-keys \
--from-literal=keys="3+ rfc4106(gcm(aes)) $(dd if=/dev/urandom count=20 bs=1 2>/dev/null | xxd -p -c 64) 128"
# ztunnel mTLS (beta, 1.19+): enable, then enroll namespaces
helm upgrade cilium cilium/cilium --version 1.20.2 --namespace kube-system \
-f cilium-values.yaml --set encryption.enabled=true --set encryption.type=ztunnel
kubectl label namespace <namespace> io.cilium/mtls-enabled=true
IPsec key format
The command matches the Cilium 1.20 IPsec guide. The + after the SPI is strongly recommended because it turns on per-tunnel keys. Some older algorithms were found insecure and deprecated in 1.16 (GHSA-pwqm-x5x6-5586).
Enable netkit (Kernel 6.8+)¶
# New nodes / new clusters only: veth and netkit cannot be mixed on one node
helm upgrade cilium cilium/cilium --version 1.20.2 --namespace kube-system \
-f cilium-values.yaml \
--set bpf.datapathMode=netkit \
--set bpf.masquerade=true \
--set kubeProxyReplacement=true
# 1.20+: bpf.datapathMode=auto picks netkit when supported, else veth
kubectl -n kube-system exec ds/cilium -- cilium-dbg status | grep -i "device mode"
netkit needs eBPF host routing (do not set bpf.hostLegacyRouting=true) and does not work with bpf.tproxy=true. For existing clusters, use per-node config on new nodes or cordon and drain nodes before switching.
Tetragon Runtime Security¶
helm install tetragon cilium/tetragon --namespace kube-system
kubectl -n kube-system rollout status ds/tetragon
# Stream events (process exec/exit plus policy events)
kubectl exec -n kube-system ds/tetragon -c tetragon -- tetra getevents -o compact
Kill any process that calls setuid(0) (privilege escalation):
apiVersion: cilium.io/v1alpha1
kind: TracingPolicy
metadata:
name: detect-priv-escalation
spec:
kprobes:
- call: sys_setuid
syscall: true
args:
- index: 0
type: int
selectors:
- matchArgs:
- index: 0
operator: Equal
values: ["0"]
matchActions:
- action: Sigkill
Kill any process that opens /etc/shadow, using the fd_install kernel function:
apiVersion: cilium.io/v1alpha1
kind: TracingPolicy
metadata:
name: monitor-sensitive-files
spec:
kprobes:
- call: "fd_install"
syscall: false
args:
- index: 0
type: "int"
- index: 1
type: "file"
selectors:
- matchArgs:
- index: 1
operator: "Equal"
values:
- "/etc/shadow"
matchActions:
- action: Sigkill
Test enforcement in audit mode first
Start without matchActions (events only), confirm the matches in tetra getevents, then add Sigkill. With syscall: true, Tetragon adds the architecture prefix (for example __x64_) to syscall names.
Troubleshooting¶
| Symptom | Diagnosis | Fix |
|---|---|---|
| Pod connectivity fails | cilium status, cilium connectivity test, hubble observe --verdict DROPPED |
Check drop reason. Verify routing mode and firewall ports (8472/udp, 6081/udp, 4240/tcp) |
| Policy not enforced | cilium-dbg endpoint list (policy enforcement column), check labels |
Verify label selectors match. Check policyEnforcementMode |
| Unexpected drops after first policy | Endpoint switched to default deny | Add the missing allow (often DNS), or use enableDefaultDeny: false for audit-style policies |
toFQDNs rule never matches |
DNS not proxied | Add an L7 dns rule for port 53 to kube-dns |
| High CPU on agent | cilium-dbg metrics list, endpoint regeneration counts |
Reduce policy churn. Check identity cardinality (avoid high-churn labels) |
| Policy map full | cilium-dbg policy map pressure commands (1.20), bpf.policyMapMax |
Simplify selectors, raise bpf.policyMapMax |
| Hubble flows missing | hubble status, cilium status |
Enable Hubble Relay. Check port 4244 between nodes and 4245 to Relay |
| DNS resolution issues | kubectl -n kube-system exec ds/cilium -- cilium-dbg monitor --type l7 |
Check DNS proxy and CoreDNS connectivity. On Alpine/musl see the DNS proxy "Refused" caveat |
| Gateway times out | Gateway Service has an address but Envoy gets no traffic | Nodes lack iptables TPROXY modules. Load them or try bpf.tproxy=true (beta) |
| Agent will not start after 1.20 upgrade | Agent logs mention Kafka/l7proto or v2alpha1 |
Remove Kafka rules. Move CiliumNodeConfig to cilium.io/v2 |
# Agent status (inside the pod the binary is cilium-dbg)
cilium status
kubectl -n kube-system exec ds/cilium -- cilium-dbg status --verbose
# Endpoints and identities
kubectl -n kube-system exec ds/cilium -- cilium-dbg endpoint list
kubectl -n kube-system exec ds/cilium -- cilium-dbg endpoint get <endpoint-id>
kubectl -n kube-system exec ds/cilium -- cilium-dbg identity list
# BPF state
kubectl -n kube-system exec ds/cilium -- cilium-dbg bpf ct list global
kubectl -n kube-system exec ds/cilium -- cilium-dbg service list
# Live drops on a node
kubectl -n kube-system exec ds/cilium -- cilium-dbg monitor --type drop
# Collect a sysdump for bug reports
cilium sysdump