Skip to content

How-to Guides

Scope

Task recipes for Cilium 1.20: install and upgrade, kube-proxy replacement and migration, network policy, Hubble, Gateway API, Cluster Mesh, encryption, netkit, Tetragon, and troubleshooting. Versions, Helm defaults and ports are in Reference. Internals are in Explanation.

Pin versions

Commands below pin 1.20.2, the latest stable release on 2026-09-25. Check https://raw.githubusercontent.com/cilium/cilium/main/stable.txt for newer patches. Only upgrade one minor version at a time.

Deployment Patterns

Installation Methods

Method Use Case Notes
Helm (https://helm.cilium.io/) Production Most flexible. Recommended for GitOps
Helm OCI (oci://quay.io/cilium/charts/cilium) Production, air-gapped mirrors Since 1.19. Charts signed with cosign
cilium CLI (cilium install) Dev, quick setup Wraps the same Helm chart and auto-detects the environment
Managed (GKE Dataplane V2, AKS "Azure CNI Powered by Cilium") Managed Kubernetes The provider manages the Cilium version and feature set
Standalone Tetragon Runtime security without Cilium CNI Separate Helm chart cilium/tetragon

Production Helm Install (kube-proxy-free)

helm repo add cilium https://helm.cilium.io/
helm repo update

# API server address is required when kube-proxy is not present
API_SERVER_IP=<control-plane-ip-or-lb>
API_SERVER_PORT=6443

helm install cilium cilium/cilium --version 1.20.2 \
  --namespace kube-system \
  --set kubeProxyReplacement=true \
  --set k8sServiceHost=${API_SERVER_IP} \
  --set k8sServicePort=${API_SERVER_PORT} \
  --set hubble.relay.enabled=true \
  --set hubble.ui.enabled=true \
  --set bpf.masquerade=true \
  --set ipam.mode=kubernetes

ipam.mode=kubernetes only makes sense when kube-controller-manager runs with --allocate-node-cidrs. Otherwise leave the Helm default (cluster-pool) and set ipam.operator.clusterPoolIPv4PodCIDRList.

The same chart from the OCI registry, with signature verification:

cosign verify \
  --certificate-identity-regexp='https://github.com/cilium/cilium/.*' \
  --certificate-oidc-issuer=https://token.actions.githubusercontent.com \
  quay.io/cilium/charts/cilium:1.20.2

helm install cilium oci://quay.io/cilium/charts/cilium --version 1.20.2 \
  --namespace kube-system --set kubeProxyReplacement=true \
  --set k8sServiceHost=${API_SERVER_IP} --set k8sServicePort=${API_SERVER_PORT}

The cosign verify identity and issuer values are the ones given in the Cilium Helm install docs ("OCI Registry" section, v1.20).

kube-proxy Replacement

Cilium can replace kube-proxy and handle service load balancing in eBPF. Check the result on a node:

kubectl -n kube-system exec ds/cilium -- cilium-dbg status | grep KubeProxyReplacement
# KubeProxyReplacement:   True   [eth0 (Direct Routing)]

kubectl -n kube-system exec ds/cilium -- cilium-dbg status --verbose | grep -A20 "KubeProxyReplacement Details"
kubectl -n kube-system exec ds/cilium -- cilium-dbg service list

On a new kubeadm cluster, skip kube-proxy at init with kubeadm init --skip-phases=addon/kube-proxy.

Migrate an Existing Cluster from kube-proxy

Not an official one-click procedure

The Cilium docs describe per-node configuration (CiliumNodeConfig) and a per-node CNI migration flow, but no step-by-step kube-proxy-to-KPR migration. The pattern below combines those documented pieces. Rehearse it on a staging cluster. The alternative is a maintenance window: set kubeProxyReplacement=true, delete kube-proxy, and restart Cilium on all nodes.

  1. Run Cilium with kubeProxyReplacement=false next to kube-proxy, and set k8sServiceHost/k8sServicePort.
  2. Create a CiliumNodeConfig (cilium.io/v2) that sets kube-proxy-replacement: "true" for nodes carrying a migration label:

    apiVersion: cilium.io/v2
    kind: CiliumNodeConfig
    metadata:
      namespace: kube-system
      name: kpr-migration
    spec:
      nodeSelector:
        matchLabels:
          io.cilium.migration/kube-proxy-replacement: "true"
      defaults:
        kube-proxy-replacement: "true"
    
  3. Add a node affinity to the kube-proxy DaemonSet so it does not run on nodes with that label.

  4. For each node: kubectl cordon, kubectl label node $NODE io.cilium.migration/kube-proxy-replacement=true, delete the Cilium pod on that node so it restarts with the new config, flush kube-proxy's iptables rules on the node (or reboot it), check with cilium-dbg status, then kubectl uncordon.
  5. When all nodes are done, set kubeProxyReplacement=true in Helm values, delete the kube-proxy DaemonSet and the CiliumNodeConfig, and remove the labels.

Upgrade Procedure

# 1. Read the upgrade notes for the target minor (1.20: Kafka rules removed,
#    TLSRoute v1, CiliumNodeConfig v2alpha1 removed)
# 2. Move to the latest patch of the current minor first, then one minor up

# Pre-flight: pre-pull images so the upgrade does not stall on pulls
helm install cilium-preflight cilium/cilium --version 1.20.2 \
  --namespace kube-system --set preflight.enabled=true \
  --set agent=false --set operator.enabled=false
kubectl -n kube-system get daemonset cilium-pre-flight-check
helm delete cilium-preflight --namespace kube-system

# Upgrade, keeping your values under version control
helm upgrade cilium cilium/cilium --version 1.20.2 \
  --namespace kube-system -f cilium-values.yaml

# Post-upgrade validation
cilium status --wait
cilium connectivity test

Avoid --reuse-values across minors

--reuse-values carries old values forward and skips new chart defaults, which often causes trouble across minor versions. Keep an explicit values file instead. L7/proxied connections (L7 policy, Ingress, Gateway API) are reset during the upgrade. L3/L4 traffic is not.

Commands & Recipes

Install the CLIs

# cilium CLI
CILIUM_CLI_VERSION=$(curl -s https://raw.githubusercontent.com/cilium/cilium-cli/main/stable.txt)
CLI_ARCH=amd64
curl -L --fail --remote-name-all \
  https://github.com/cilium/cilium-cli/releases/download/${CILIUM_CLI_VERSION}/cilium-linux-${CLI_ARCH}.tar.gz{,.sha256sum}
sha256sum --check cilium-linux-${CLI_ARCH}.tar.gz.sha256sum
sudo tar xzvfC cilium-linux-${CLI_ARCH}.tar.gz /usr/local/bin

# Hubble CLI
HUBBLE_VERSION=$(curl -s https://raw.githubusercontent.com/cilium/hubble/main/stable.txt)
curl -L --fail --remote-name-all \
  https://github.com/cilium/hubble/releases/download/$HUBBLE_VERSION/hubble-linux-amd64.tar.gz{,.sha256sum}
sha256sum --check hubble-linux-amd64.tar.gz.sha256sum
sudo tar xzvfC hubble-linux-amd64.tar.gz /usr/local/bin

Quick Install and Validate

# Quick install with the CLI (auto-detects cluster type)
cilium install --version 1.20.2
cilium status --wait
cilium connectivity test

# Or with Helm and Hubble enabled
helm install cilium cilium/cilium --version 1.20.2 \
  --namespace kube-system \
  --set hubble.enabled=true \
  --set hubble.relay.enabled=true \
  --set hubble.ui.enabled=true

Hubble Observability

# Enable Relay and UI on an existing install
cilium hubble enable --ui

# Port-forward Hubble Relay (localhost:4245)
cilium hubble port-forward &

# Observe flows
hubble observe --namespace default --follow
hubble observe --verdict DROPPED --follow
hubble observe --pod default/myapp --protocol http
hubble observe --to-fqdn "*.amazonaws.com"
hubble observe --from-pod default/frontend --to-pod default/backend --verdict DROPPED

# Service map in the UI
cilium hubble ui          # or: kubectl -n kube-system port-forward svc/hubble-ui 12000:80

Network Policies

L7 HTTP policy that only allows GET /api/* from frontend to backend:

apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: allow-api-get
spec:
  endpointSelector:
    matchLabels:
      app: backend
  ingress:
    - fromEndpoints:
        - matchLabels:
            app: frontend
      toPorts:
        - ports:
            - port: "8080"
              protocol: TCP
          rules:
            http:
              - method: GET
                path: "/api/.*"

DNS-based egress policy. toFQDNs only works if DNS traffic also goes through the DNS proxy, which is what the second rule does:

apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: allow-googleapis
spec:
  endpointSelector:
    matchLabels:
      app: myapp
  egress:
    - toEndpoints:
        - matchLabels:
            io.kubernetes.pod.namespace: kube-system
            k8s-app: kube-dns
      toPorts:
        - ports:
            - port: "53"
              protocol: ANY
          rules:
            dns:
              - matchPattern: "*"
    - toFQDNs:
        - matchPattern: "**.googleapis.com"   # multi-level wildcard, 1.19+
      toPorts:
        - ports:
            - port: "443"
              protocol: TCP

Kafka rules removed in 1.20

Policies with rules.kafka, rules.l7 or rules.l7proto must be cleaned up before upgrading to 1.20. Find them with: kubectl get cnp,ccnp -A -o json | jq -r '.items[] | select(.. | objects | has("kafka") or has("l7proto")) | "\(.metadata.namespace)/\(.metadata.name)"'

Default Deny Baseline

Selecting endpoints with a rule that has an empty ingress/egress section switches them to default deny. This namespace-scoped example denies all egress for role: restricted pods (from the Cilium examples):

apiVersion: "cilium.io/v2"
kind: CiliumNetworkPolicy
metadata:
  name: "deny-all-egress"
spec:
  endpointSelector:
    matchLabels:
      role: restricted
  egress:
  - {}

A cluster-wide baseline usually puts all workload pods (outside kube-system) into default deny and still allows DNS, so later namespace policies only add allows:

apiVersion: cilium.io/v2
kind: CiliumClusterwideNetworkPolicy
metadata:
  name: baseline-default-deny
spec:
  endpointSelector:
    matchExpressions:
      - key: io.kubernetes.pod.namespace
        operator: NotIn
        values: ["kube-system"]
  ingress:
    - {}
  egress:
    - toEndpoints:
        - matchLabels:
            io.kubernetes.pod.namespace: kube-system
            k8s-app: kube-dns
      toPorts:
        - ports:
            - port: "53"
              protocol: ANY
          rules:
            dns:
              - matchPattern: "*"

Do not use deny-all against all

A policy with ingressDeny: [{fromEntities: ["all"]}] and egressDeny: [{toEntities: ["all"]}] blocks everything, and no allow rule can override it, because deny always wins. Use deny rules only for targeted blocks (for example ingress from world).

Gateway API

# 1. Install Gateway API v1.6.1 CRDs (standard channel); add tcproutes/udproutes/listenersets if needed
GW=https://raw.githubusercontent.com/kubernetes-sigs/gateway-api/v1.6.1/config/crd/standard
for crd in gatewayclasses gateways httproutes referencegrants grpcroutes backendtlspolicies tlsroutes; do
  kubectl apply --server-side -f ${GW}/gateway.networking.k8s.io_${crd}.yaml
done

# 2. Enable the controller (requires kube-proxy replacement and l7Proxy)
helm upgrade cilium cilium/cilium --version 1.20.2 --namespace kube-system \
  -f cilium-values.yaml --set gatewayAPI.enabled=true
kubectl -n kube-system rollout restart deployment/cilium-operator
kubectl -n kube-system rollout restart ds/cilium
kubectl get gatewayclass cilium

Existing TLSRoutes

If you already use TLSRoute v1alpha2, back it up and install the experimental v1.6.1 TLSRoute CRD instead of the standard one. See Reference: Gateway API Support.

Cluster Mesh

# Each cluster needs a unique name and ID (and non-overlapping pod CIDRs)
cilium install --set cluster.name=$CLUSTER1 --set cluster.id=1 --context $CLUSTER1
cilium install --set cluster.name=$CLUSTER2 --set cluster.id=2 --context $CLUSTER2

cilium clustermesh enable --context $CLUSTER1
cilium clustermesh enable --context $CLUSTER2
cilium clustermesh status --context $CLUSTER1 --wait

cilium clustermesh connect --context $CLUSTER1 --destination-context $CLUSTER2
cilium clustermesh status --context $CLUSTER1 --wait
cilium connectivity test --context $CLUSTER1 --multi-cluster $CLUSTER2

Mark a Service global by annotating it service.cilium.io/global: "true" in each cluster. With Helm-generated certificates, Cluster Mesh certificates expire after one year by default (1.20). Re-render the chart (upgrade) at least yearly, or use the cronJob/certmanager generation methods.

Enable Transparent Encryption

# WireGuard (UDP 51871 must be open between nodes)
helm upgrade cilium cilium/cilium --version 1.20.2 --namespace kube-system \
  -f cilium-values.yaml \
  --set encryption.enabled=true \
  --set encryption.type=wireguard
kubectl -n kube-system exec ds/cilium -- cilium-dbg encrypt status

# IPsec: create the pre-shared key first (or: cilium encrypt create-key --auth-algo rfc4106-gcm-aes)
kubectl create -n kube-system secret generic cilium-ipsec-keys \
  --from-literal=keys="3+ rfc4106(gcm(aes)) $(dd if=/dev/urandom count=20 bs=1 2>/dev/null | xxd -p -c 64) 128"

# ztunnel mTLS (beta, 1.19+): enable, then enroll namespaces
helm upgrade cilium cilium/cilium --version 1.20.2 --namespace kube-system \
  -f cilium-values.yaml --set encryption.enabled=true --set encryption.type=ztunnel
kubectl label namespace <namespace> io.cilium/mtls-enabled=true

IPsec key format

The command matches the Cilium 1.20 IPsec guide. The + after the SPI is strongly recommended because it turns on per-tunnel keys. Some older algorithms were found insecure and deprecated in 1.16 (GHSA-pwqm-x5x6-5586).

Enable netkit (Kernel 6.8+)

# New nodes / new clusters only: veth and netkit cannot be mixed on one node
helm upgrade cilium cilium/cilium --version 1.20.2 --namespace kube-system \
  -f cilium-values.yaml \
  --set bpf.datapathMode=netkit \
  --set bpf.masquerade=true \
  --set kubeProxyReplacement=true
# 1.20+: bpf.datapathMode=auto picks netkit when supported, else veth
kubectl -n kube-system exec ds/cilium -- cilium-dbg status | grep -i "device mode"

netkit needs eBPF host routing (do not set bpf.hostLegacyRouting=true) and does not work with bpf.tproxy=true. For existing clusters, use per-node config on new nodes or cordon and drain nodes before switching.

Tetragon Runtime Security

helm install tetragon cilium/tetragon --namespace kube-system
kubectl -n kube-system rollout status ds/tetragon
# Stream events (process exec/exit plus policy events)
kubectl exec -n kube-system ds/tetragon -c tetragon -- tetra getevents -o compact

Kill any process that calls setuid(0) (privilege escalation):

apiVersion: cilium.io/v1alpha1
kind: TracingPolicy
metadata:
  name: detect-priv-escalation
spec:
  kprobes:
    - call: sys_setuid
      syscall: true
      args:
        - index: 0
          type: int
      selectors:
        - matchArgs:
            - index: 0
              operator: Equal
              values: ["0"]
          matchActions:
            - action: Sigkill

Kill any process that opens /etc/shadow, using the fd_install kernel function:

apiVersion: cilium.io/v1alpha1
kind: TracingPolicy
metadata:
  name: monitor-sensitive-files
spec:
  kprobes:
    - call: "fd_install"
      syscall: false
      args:
        - index: 0
          type: "int"
        - index: 1
          type: "file"
      selectors:
        - matchArgs:
            - index: 1
              operator: "Equal"
              values:
                - "/etc/shadow"
          matchActions:
            - action: Sigkill

Test enforcement in audit mode first

Start without matchActions (events only), confirm the matches in tetra getevents, then add Sigkill. With syscall: true, Tetragon adds the architecture prefix (for example __x64_) to syscall names.

Troubleshooting

Symptom Diagnosis Fix
Pod connectivity fails cilium status, cilium connectivity test, hubble observe --verdict DROPPED Check drop reason. Verify routing mode and firewall ports (8472/udp, 6081/udp, 4240/tcp)
Policy not enforced cilium-dbg endpoint list (policy enforcement column), check labels Verify label selectors match. Check policyEnforcementMode
Unexpected drops after first policy Endpoint switched to default deny Add the missing allow (often DNS), or use enableDefaultDeny: false for audit-style policies
toFQDNs rule never matches DNS not proxied Add an L7 dns rule for port 53 to kube-dns
High CPU on agent cilium-dbg metrics list, endpoint regeneration counts Reduce policy churn. Check identity cardinality (avoid high-churn labels)
Policy map full cilium-dbg policy map pressure commands (1.20), bpf.policyMapMax Simplify selectors, raise bpf.policyMapMax
Hubble flows missing hubble status, cilium status Enable Hubble Relay. Check port 4244 between nodes and 4245 to Relay
DNS resolution issues kubectl -n kube-system exec ds/cilium -- cilium-dbg monitor --type l7 Check DNS proxy and CoreDNS connectivity. On Alpine/musl see the DNS proxy "Refused" caveat
Gateway times out Gateway Service has an address but Envoy gets no traffic Nodes lack iptables TPROXY modules. Load them or try bpf.tproxy=true (beta)
Agent will not start after 1.20 upgrade Agent logs mention Kafka/l7proto or v2alpha1 Remove Kafka rules. Move CiliumNodeConfig to cilium.io/v2
# Agent status (inside the pod the binary is cilium-dbg)
cilium status
kubectl -n kube-system exec ds/cilium -- cilium-dbg status --verbose

# Endpoints and identities
kubectl -n kube-system exec ds/cilium -- cilium-dbg endpoint list
kubectl -n kube-system exec ds/cilium -- cilium-dbg endpoint get <endpoint-id>
kubectl -n kube-system exec ds/cilium -- cilium-dbg identity list

# BPF state
kubectl -n kube-system exec ds/cilium -- cilium-dbg bpf ct list global
kubectl -n kube-system exec ds/cilium -- cilium-dbg service list

# Live drops on a node
kubectl -n kube-system exec ds/cilium -- cilium-dbg monitor --type drop

# Collect a sysdump for bug reports
cilium sysdump

Sources