Linkerd Explanation¶
What this page covers
How Linkerd works and why it is built the way it is: the control plane, the Rust micro-proxy, traffic interception, identity and mTLS, authorization policy, load balancing and circuit breaking, the move from ServiceProfiles to the Gateway API, multicluster, egress, native sidecars, and the 2024 release-model change. Exact values (ports, defaults, CRD versions, benchmark numbers) are in the Reference. Tasks are in the How-to Guides.
Overview¶
Linkerd is a security-first service mesh for Kubernetes. It has two parts:
- The control plane is a small set of deployments in the
linkerdnamespace. - The data plane is one
linkerd2-proxyper meshed pod.linkerd2-proxyis a Rust micro-proxy written for this single job. It is not a general-purpose proxy.
mTLS between meshed pods is on by default and needs no configuration. The project's stated design principle is operational simplicity: features should work without configuration where possible, and need minimal, consistent configuration where not.
Linkerd 2.x was a rewrite (Go control plane, Rust proxy) of the original JVM-based Linkerd 1.x. The CNCF graduated it on 2021-07-28, making it the first service mesh to graduate. The current major version is 2.20 (2026-06-23).
Architecture at a Glance¶
The diagram below shows the core control plane, the optional extensions, and two meshed pods. Proxies pull discovery, policy and certificates from the control plane and talk to each other over mTLS.
flowchart TB
subgraph CP["Control plane (namespace linkerd)"]
subgraph DESTPOD["linkerd-destination pod"]
DEST["destination<br/>(Go: discovery, profiles)"]
POLICY["policy controller<br/>(Rust: policy API, webhook)"]
SPV["sp-validator<br/>(ServiceProfile webhook)"]
end
IDENT["linkerd-identity<br/>(CA, signs CSRs)"]
INJ["linkerd-proxy-injector<br/>(mutating webhook)"]
HB["linkerd-heartbeat<br/>(CronJob)"]
end
subgraph EXT["Optional extensions"]
VIZ["linkerd-viz<br/>(web, metrics-api, tap, Prometheus)"]
MC["linkerd-multicluster<br/>(service mirror, linkerd-gateway)"]
CNI["linkerd-cni<br/>(DaemonSet, replaces linkerd-init)"]
end
KAPI["Kubernetes API<br/>(Services, EndpointSlices, Pods, CRDs)"]
subgraph PODA["Meshed pod A"]
APPA["app container"]
PA["linkerd-proxy<br/>(native sidecar)"]
end
subgraph PODB["Meshed pod B"]
PB["linkerd-proxy<br/>(native sidecar)"]
APPB["app container"]
end
KAPI -->|watch| DEST
KAPI -->|watch policy CRDs| POLICY
KAPI -->|"pod CREATE admission"| INJ
INJ -.->|"adds linkerd-init + linkerd-proxy"| PODA
PA -->|"gRPC :8086 discovery"| DEST
PA -->|"gRPC :8090 policy"| POLICY
PA -->|"CSR :8080"| IDENT
APPA --> PA
PA <-->|"mTLS (TLS 1.3, ML-KEM-768 hybrid)"| PB
PB --> APPB
VIZ -.->|"scrape :4191/metrics"| PA
Control Plane Components¶
Destination¶
The destination service is how proxies learn about the world. It:
- Watches Kubernetes Services, EndpointSlices, Pods, ServiceProfiles and Gateway API routes.
- Answers proxy lookups over the destination gRPC API (
linkerd2-proxy-api) on port 8086. An answer holds the endpoint set, the TLS identity expected on the other end, protocol hints (for example opaque ports orappProtocol), and zone metadata. - Serves route configuration: HTTPRoute and GRPCRoute rules, retries, timeouts, and ServiceProfile routes for older setups.
- Holds the model of the cluster, so it uses most of the control plane's memory. Linkerd 2.20 refactored its internal state management and reports memory cuts "in some cases by almost 85%" on large clusters with high pod churn.
The same pod runs two more containers:
- Policy controller (
policy, written in Rust). It serves inbound and outbound policy (Servers, AuthorizationPolicies, routes, rate limits, EgressNetworks) on port 8090. It also runs the admission webhook that validates policy resources. - sp-validator. It is the admission webhook that validates
ServiceProfileresources. It does not validate Server or policy CRDs.
Identity¶
The identity service is the mesh's built-in certificate authority (CA):
- It accepts certificate signing requests (CSRs) from proxies on port 8080. Each CSR includes the pod's bound ServiceAccount token. The identity service checks the token with the Kubernetes API before it signs.
- It signs with the issuer certificate and key, stored in a Secret in the
linkerdnamespace that only the identity ServiceAccount can read. - It issues leaf certificates that are valid for 24 hours by default. Proxies renew them automatically.
- The identity name is
<serviceaccount>.<namespace>.serviceaccount.identity.linkerd.cluster.local.
Proxy Injector¶
The proxy injector is a mutating admission webhook. It gets called on every pod creation and acts when linkerd.io/inject: enabled (or ingress) is set on the pod or its namespace. It adds:
linkerd-init, an init container that installsiptablesrules. It is left out when thelinkerd-cniplugin is used.linkerd-proxy. Since Linkerd 2.20 this is added by default as a native sidecar: an init container withrestartPolicy: Always.
It also applies per-workload overrides from config.linkerd.io/* annotations. The injector skips kube-system and cert-manager so that it never blocks the components Linkerd itself depends on. Adding the annotation does not change pods that already exist. They must be restarted.
Heartbeat and Extensions¶
linkerd-heartbeat is a daily CronJob that sends anonymous version and usage data. Disable it with disableHeartBeat: true. Everything else is an extension with its own namespace and lifecycle:
- viz: dashboard,
tap,metrics-api, optional Prometheus. - multicluster: service mirror and gateway.
- jaeger: tracing.
- smi: TrafficSplit. Deprecated.
The dashboard is not part of the core control plane. The data plane does not need any extension to run.
Data Plane: linkerd2-proxy¶
The following sequence shows how a pod joins the mesh, gets its identity, and sends its first request.
sequenceDiagram
participant K8s as Kubernetes API
participant Inj as proxy-injector
participant Init as linkerd-init
participant Proxy as linkerd-proxy
participant Id as identity
participant Dest as destination + policy
participant App as app container
K8s->>Inj: AdmissionReview (pod CREATE, inject=enabled)
Inj-->>K8s: JSON patch adds linkerd-init + linkerd-proxy
Init->>Init: iptables REDIRECT to 4143 (in) and 4140 (out)
Proxy->>Proxy: Generate private key in tmpfs
Proxy->>Id: CSR + ServiceAccount token (validated with trust anchor)
Id-->>Proxy: Leaf cert for sa.ns.serviceaccount.identity.linkerd.cluster.local
Proxy-->>App: Ready (native sidecar starts before app)
App->>Proxy: connect(service ClusterIP), redirected to 4140
Proxy->>Dest: Resolve ClusterIP (endpoints, expected identity, routes, policy)
Dest-->>Proxy: Stream of updates
Proxy->>Proxy: Pick endpoint (P2C over EWMA latency)
Proxy Design¶
The proxy is built on the Rust async ecosystem (Tokio, Hyper, Tower). It provides:
- Protocol detection: it recognizes HTTP/1.x, HTTP/2 and gRPC. Anything else is proxied as opaque TCP. Since 2.18 it can skip detection when a Service port declares
appProtocol. - L7 features on HTTP and gRPC: per-request load balancing, retries, timeouts, circuit breaking, dynamic routing through HTTPRoute and GRPCRoute, and per-route metrics.
- L4: connection-level load balancing for TCP.
- Transparent mTLS between meshed pods.
- Telemetry: Prometheus metrics on the admin port
:4191. Prometheus scrapes them (for example the one in linkerd-viz). Proxies do not push metrics to the control plane. Since 2.17 the proxy also exports OpenTelemetry traces. - HTTP/1.1 to HTTP/2 upgrade between proxies (
enableH2Upgrade, on by default). This multiplexes HTTP/1 requests over one mTLS connection between two proxies. - Tap API on
:4190for live request inspection. Access is controlled by RBAC.
Why Rust and not Envoy? Linkerd's maintainers argue that a purpose-built proxy in a memory-safe language gives a smaller footprint and attack surface, and that it removes most configuration surface: there is no xDS, no filter chains, and no Wasm. Buoyant's 2021 benchmark measured a maximum proxy memory of 17.8 MB against 154.6 MB for Istio's Envoy sidecar (see the Reference). The cost is extensibility. You cannot add custom filters, and features arrive only through Linkerd releases. Since 2.19, the proxy's TLS stack uses aws-lc instead of ring. That change enabled post-quantum key exchange and FIPS builds in BEL.
Meshed Connections: Outbound vs Inbound¶
The proxy in the pod that opens a connection is the outbound proxy. The proxy in the pod that accepts it is the inbound proxy. They do different jobs:
| Side | Responsibilities |
|---|---|
| Outbound (client) | Service discovery, load balancing, circuit breaking, retries, timeouts, egress policy, route metrics |
| Inbound (server) | Authorization policy, rate limiting, inbound metrics (much richer since 2.20) |
This split explains several rules. The client must be meshed for retries, timeouts or circuit breaking to work. The server must be meshed for authorization to be enforced.
Transparent Interception¶
linkerd-init (or linkerd-cni) programs iptables NAT rules inside the pod's network namespace.
- Inbound:
PREROUTINGjumps toPROXY_INIT_REDIRECT. Ports inskip-inbound-portsare left alone. All other TCP is sent withREDIRECTto port 4143. - Outbound:
OUTPUTjumps toPROXY_INIT_OUTPUT. The chain skips packets owned by the proxy UID 2102 (so proxy traffic is not captured again), traffic on loopback, andskip-outbound-ports. Everything else goes withREDIRECTto port 4140.
REDIRECT keeps the original destination. The proxy reads it with the SO_ORIGINAL_DST socket option, then asks the destination service what that IP:port is: a Service ClusterIP, a pod IP, or something outside the cluster.
Linkerd depends on the ClusterIP being present on outbound packets. If Cilium's kube-proxy replacement does socket-level load balancing inside pods, the proxy only sees a pod IP. mTLS and metrics still work, but EWMA balancing and HTTPRoute routing do not. The docs recommend socketLB.hostNamespaceOnly=true for Cilium.
linkerd-init needs CAP_NET_ADMIN. The CNI plugin moves rule setup into a chained CNI plugin on each node, so application pods need no elevated capabilities. linkerd-init also supports an nft (nftables) mode.
Protocol Detection and Opaque Ports¶
Protocol detection lets you "drop in" Linkerd without configuration. The proxy peeks at the first bytes of a connection to decide between HTTP and opaque TCP. Two cases break this:
- Server-speaks-first protocols (MySQL, SMTP and others). The client waits for the server, so there is nothing to peek at. Detection waits for a timeout and then falls back to TCP. These ports must be marked opaque. The chart's default opaque list is
25, 587, 3306, 4444, 5432, 6379, 9300, 11211. - Extreme load. The application may not send data before the detection timeout. The connection is silently downgraded to TCP and loses its HTTP features.
Linkerd 2.18 added protocol declaration to fix the second case. If a Service port sets appProtocol (for example http), the proxies skip detection and pass the declared protocol between each other.
Identity, mTLS and the Trust Chain¶
Linkerd uses a three-level PKI:
flowchart TB
TA["Trust anchor (root CA)<br/>ConfigMap linkerd-identity-trust-roots<br/>default 365 days if CLI-generated"]
ISS["Issuer cert + key (intermediate CA)<br/>Secret linkerd-identity-issuer<br/>default 1 year if CLI-generated"]
LEAF["Proxy leaf cert<br/>sa.ns.serviceaccount.identity.linkerd.cluster.local<br/>24 h, auto-renewed"]
TA -->|signs| ISS
ISS -->|"identity service signs CSR"| LEAF
CM["cert-manager + trust-manager<br/>(optional automation)"] -.->|"rotates issuer, distributes bundle"| ISS
How a meshed connection gets authenticated:
- The proxy receives the trust anchor in an environment variable at injection time. It generates a private key in a tmpfs
emptyDir, so the key stays in memory and never leaves the pod. - It sends a CSR plus its bound ServiceAccount token to
identity. It validates the identity endpoint against the trust anchor. It gets back a leaf certificate that works as both client and server certificate. - When the app connects to a meshed destination, the destination service tells the client proxy which identity to expect. The client proxy checks that the server certificate chains to the trust anchor and carries that identity.
- The server (inbound) proxy learns the client's identity from the client certificate. It uses that identity in authorization policy (for example
MeshTLSAuthentication).
For all meshed traffic, the TLS parameters are TLS 1.3 and hybrid ML-KEM-768 + X25519 key exchange. That makes the key exchange post-quantum, and it is on by default since 2.19. The docs list the AES_128_GCM cipher suite. The 2.19 announcement says AES_256_GCM support was added. The TLS algorithms in use are exported as metrics.
Operational catch: certificate expiry
The CLI-generated trust anchor and issuer both expire after one year. When they expire, proxies can no longer get certificates, and meshed traffic fails. linkerd check warns as expiry approaches. For production, bring your own trust anchor with a long lifetime and let cert-manager rotate the issuer (see How-to Guides). BEL 2.20 adds automated trust anchor rotation as an enterprise feature.
Clusters that are linked for multicluster must share a trust anchor. The default linkerd install credentials will not work for that.
SPIFFE. Kubernetes workloads get the DNS-style identity above. Since 2.15, workloads outside Kubernetes (mesh expansion) get SPIFFE IDs from SPIRE. Authorization policy can use both kinds of identity.
Caveats of "mTLS by default"¶
- mTLS covers meshed-to-meshed TCP only. Traffic from unmeshed clients (for example kubelet probes) is plaintext.
- By default, an inbound proxy accepts plaintext from unmeshed sources. The default inbound policy is
all-unauthenticated. Useall-authenticated,deny, or policy CRDs to require mTLS. - Ports in
skip-inbound-portsorskip-outbound-portsbypass the proxy entirely. They get no encryption and no metrics.
Authorization Policy¶
Linkerd's zero-trust model: each pod makes its own authorization decisions, holds only its own keys, and decides based on cryptographic workload identity, not IP addresses. The maintainers argue that this per-pod boundary is a security advantage of sidecars over shared node proxies.
Policy has two layers:
- Default policy. It is set cluster-wide at install and can be overridden with an annotation on the namespace or pod. It is fixed when the proxy starts.
- Dynamic CRDs. They update on the fly:
- Targets:
Server(a port on a set of pods),HTTPRouteandGRPCRoute(subsets of requests). - Authentications:
MeshTLSAuthentication(identities or ServiceAccounts) andNetworkAuthentication(CIDRs). AuthorizationPolicyties authentications to targets.
- Targets:
The older ServerAuthorization (targets Server only) still works but is superseded.
The following flowchart shows how an inbound proxy decides whether to admit a request.
flowchart TD
REQ["Inbound connection or request<br/>at inbound proxy :4143"] --> SRV{"Server selects<br/>this pod + port?"}
SRV -->|No| DEF{"Default policy<br/>(annotation or cluster)"}
SRV -->|Yes| ROUTE{"HTTPRoute / GRPCRoute<br/>matches?"}
ROUTE -->|"No routes defined"| AP{"AuthorizationPolicy<br/>on Server satisfied?"}
ROUTE -->|Match| APR{"AuthorizationPolicy<br/>on route satisfied?"}
ROUTE -->|"Routes exist, none match"| DENY
AP -->|Yes| ALLOW["Forward to app"]
APR -->|Yes| ALLOW
AP -->|No| AM{"Server accessPolicy"}
APR -->|No| AM
AM -->|audit| AUDIT["Forward, but log + metric<br/>authz_name=audit"]
AM -->|deny| DENY["HTTP 403 or TCP refused"]
DEF -->|"allows client"| ALLOW
DEF -->|"denies client"| DENY
Design details worth knowing:
- Probes. Linkerd authorizes kubelet probes automatically, but only when a
Serverhas no routes. Once any HTTPRoute is attached to aServer, you must authorize probe paths yourself. - Audit mode (since 2.16).
accessPolicy: auditon aServer, orauditas the default policy, logs and counts violations without blocking them. This lets you roll out policy safely. - Rate limiting (since 2.17). Enforcement is server-side and local: each inbound proxy limits its own pod. It uses the Generic Cell Rate Algorithm (GCRA) with a 1-second tolerance. Limits can be set as a total, per client identity (fairness), or as overrides for specific clients. All unmeshed clients count as one client. There is no global (cluster-wide) rate limiter.
- HA mode. Policy can only be enforced if the proxy is present. HA mode sets the injector webhook's failure policy to
Fail, so annotated pods cannot start without a proxy.
Load Balancing and Circuit Breaking¶
For HTTP, HTTP/2 and gRPC, Linkerd balances each request, not each connection. This matters for gRPC, where kube-proxy's connection-level balancing pins all of a client's traffic to one pod. The algorithm is power of two choices (P2C) over peak-EWMA latency: pick two endpoints at random and send to the one with the lower exponentially weighted moving average of latency, weighted by outstanding requests.
The flowchart below shows one balancing decision, with the circuit breaker removing failing endpoints before selection.
flowchart LR
REQ["Request from app"] --> CB["Failure accrual filter<br/>(drop endpoints with tripped breakers)"]
CB --> P2C["P2C: sample 2 ready endpoints"]
P2C --> CMP{"Compare peak-EWMA cost<br/>(latency x in-flight)"}
CMP -->|"lower cost"| E1["pod-a (EWMA 5 ms)"]
CMP -.->|"higher cost"| E2["pod-b (EWMA 200 ms)"]
E1 --> OBS["Observe latency + status<br/>update EWMA and accrual"]
The EWMA update takes the form ewma = (1 - α) × ewma + α × latest_rtt, with a decay that reacts quickly to latency spikes. As a result, Linkerd routes around slow endpoints without separate health checks. For pod IPs (headless Services or direct pod addressing), Linkerd does no load balancing. For destinations outside the cluster, it balances across DNS results.
Circuit breaking has existed since 2.13 and is opt-in per Service through balancer.linkerd.io/failure-accrual:
consecutive: trip after N consecutive failures (default 7). A failure is a 5xx or one of the gRPC failure codes. After a minimum penalty, the endpoint enters probation and must succeed on a real request to return. Health checks do not count. Backoff grows exponentially with jitter up to a maximum penalty.unified(2.20, experimental): trips on success rate below a threshold (default 0.8 over 10 s) or on consecutive failures. For success rate, HTTP 429 and gRPCRESOURCE_EXHAUSTEDcount as failures.
Rate-limit-aware load balancing (2.20, experimental, the "Load Biaser"). Rate-limited responses are usually fast, so plain EWMA would send more traffic to a throttling pod. The Load Biaser replaces the latency of a 429 with a penalty (default 5 s), or with Retry-After if that is larger. This steers traffic away from the throttling pod.
ServiceProfiles disable the new machinery
If a ServiceProfile exists for a Service, proxies use its retry configuration and do not perform HTTPRoute retries or circuit breaking for it. Migrate to Gateway API routes before turning these features on.
Configuration Model: From ServiceProfiles to the Gateway API¶
| Era | Mechanism | Status (2.20) |
|---|---|---|
| 2.1 to 2.10 | ServiceProfile (linkerd.io/v1alpha2): regex routes, retry budgets, timeouts, per-route metrics |
Supported but frozen since 2.16 |
| 2.x (SMI) | TrafficSplit through the linkerd-smi extension |
Deprecated |
| 2.12 to 2.13 | Linkerd's own policy.linkerd.io HTTPRoute for policy and routing |
Supported, no new features |
| 2.14+ | Gateway API HTTPRoute / GRPCRoute (gateway.networking.k8s.io). Linkerd 2.14 was the first mesh conformant with the GAMMA Mesh profile. |
Primary interface |
| 2.16+ | Retries, timeouts and route metrics as annotations on Gateway API routes. Timeouts and retries compose. | Current |
Why the change. The Gateway API (with GAMMA for east-west traffic) gives a vendor-neutral configuration surface. Linkerd bet on it early to be "future-proof" and to avoid proprietary CRDs.
The friction. Up to 2.18, Linkerd bundled the Gateway API CRDs. Other projects then started requiring specific Gateway API versions, and the bundles conflicted. Since 2.19, the linkerd-crds chart defaults to installGatewayAPI: false. You install the Gateway API yourself and keep it within Linkerd's supported range (see the Reference).
Retries in the 2.16+ model are opt-in. By default a failed request gets one retry, and retries should only be used for idempotent requests. Per-request overrides through l5d-retry-* and l5d-timeout headers exist, but they must be enabled explicitly. Never enable them on pods that receive untrusted traffic.
Multicluster¶
Linkerd's multicluster extension mirrors Services between clusters. Cross-cluster calls use the same mTLS, metrics and policy as local ones. It supports three modes, and each Service can use a different one.
flowchart LR
subgraph WEST["Source cluster west"]
CL["client pod + proxy"]
SM["service-mirror controller<br/>(per Link)"]
MIR["podinfo-east<br/>(mirror Service)"]
FED["bb-federated<br/>(federated Service)"]
end
subgraph EAST["Target cluster east"]
GW["linkerd-gateway<br/>(LoadBalancer, proxy only)"]
SVC["podinfo<br/>label mirror.linkerd.io/exported=true"]
BBE["bb<br/>label mirror.linkerd.io/federated=member"]
end
SM -->|"watch via Link kubeconfig"| SVC
SM --> MIR
SM --> FED
CL -->|"hierarchical: mTLS via gateway"| GW
GW --> SVC
CL -->|"flat / federated: direct pod-to-pod mTLS"| BBE
| Mode | Since | Data path | Use when |
|---|---|---|---|
| Hierarchical (gateway) | 2.8 | Client proxy to the remote linkerd-gateway (LoadBalancer), then to the pod |
Clusters without pod-to-pod reachability, or across the Internet |
| Flat (pod-to-pod) | 2.14 | Direct pod-to-pod mTLS | Shared flat network. Lower latency, no LoadBalancer cost, and client identity is kept for policy. |
| Federated services | 2.17 | Direct. One logical <svc>-federated Service spans every member cluster. |
Many clusters running the same service. Clients call one name, and EWMA balances across all clusters. |
Why federated services matter: early multicluster use was pairs of "pet" clusters. Today's platforms run hundreds or thousands of clusters, all managed with GitOps. With a federated service, a client does not need to know which cluster serves a request. Failover across a down cluster, a failing service, or a slow cluster falls out of the normal latency-aware balancing. Since 2.18, Link resources and their credentials can be generated with linkerd multicluster link-gen and managed declaratively. This makes multicluster GitOps-compatible. Headless Services can be mirrored (for StatefulSets) but cannot join a federated service.
Egress Visibility and Control¶
Before 2.17, a meshed pod's traffic to the Internet was opaque to Linkerd. Linkerd 2.17 added egress handling inside the sidecar. No separate egress gateway is needed, but you can still use one.
- An
EgressNetworkresource (namespace-scoped, or global when created in the configured egress namespace) classifies traffic to destinations outside the cluster. It sets a defaulttrafficPolicyofAlloworDeny. - Metrics carry
parent_kind="EgressNetwork". With hostname labels enabled, they show the destination host, and for HTTP they show routes and status codes. - Gateway API
HTTPRoute,GRPCRoute,TLSRouteandTCPRouteresources can use anEgressNetworkas their parent. That allows allowlisting by DNS name, path or SNI, not only by IP and port. A denied HTTP request gets403.
Egress policy is not a security boundary on its own
The Linkerd docs are explicit: a workload that bypasses its sidecar (for example through skip ports, or with host networking) bypasses egress policy. Combine it with Kubernetes NetworkPolicy or CNI-level controls for real enforcement.
Native Sidecars¶
Classic sidecars caused two long-standing problems:
- Jobs never finish. The proxy keeps running after the job container exits.
- Startup races. Init containers that need the network run before the proxy exists. The app may also start before the proxy is ready. Linkerd worked around this with a
postStarthook that runslinkerd-await.
Kubernetes native sidecars (KEP-753: init containers with restartPolicy: Always) fix both. The kubelet starts them before the app and stops them after it.
| Linkerd version | Native sidecar status |
|---|---|
| 2.15 | Alpha, opt-in |
| 2.19 | Beta, opt-in through config.linkerd.io/proxy-enable-native-sidecar |
| 2.20 | Stable and on by default (proxy.nativeSidecar: true) |
You can opt out per workload or namespace with the annotation, or globally with the Helm value.
Release Model, Governance and Sustainability¶
- February 2024 (2.15). The project stopped publishing open-source stable release artifacts, including point releases and backports. Open-source users get edge releases (
edge-YY.M.N, weekly or near-weekly). Edge releases are considered production-ready unless flagged. Stable, semantically versioned builds come from vendors, in practice Buoyant Enterprise for Linkerd (BEL). The code stays Apache-2.0, and governance was not changed. - BEL licensing. BEL needs a license key. According to Buoyant, companies with fewer than 50 employees can run it free in production (without support). Larger companies pay. The list price at launch was reported as US$2,000 per cluster per month. Check Buoyant's current pricing page before relying on that figure. Some features are BEL-only: HAZL (high-availability zonal load balancing), FIPS builds, Windows workloads, policy generation, and automated trust anchor rotation in 2.20.
- October 2024. Buoyant announced that it was profitable and that Linkerd was sustainable ("Linkerd forever"). Releases since then (2.17 to 2.20) have each added several features.
- Governance. Linkerd is a CNCF Graduated project (since 2021-07-28) with a steering committee of end users. As of 2026-09, all three listed maintainers are Buoyant employees, so the project depends on one vendor.
What this means in practice
Running open-source Linkerd in production means tracking edge releases. Read each release's notes and keep the data plane within one major version of the control plane. If you need semver, backports or FIPS, you need a commercial distribution.
Performance Characteristics¶
Why Linkerd is usually lighter than an Envoy sidecar mesh:
- One-purpose proxy. There is no generic filter chain, no Wasm runtime, and no xDS configuration to hold in memory. Per-proxy state is limited to what the destination service streams for the Services the pod actually calls.
- Rust with no garbage collector. Memory is predictable and tail latency has no GC pauses.
- Per-request balancing and HTTP/2 multiplexing between proxies.
The only published head-to-head numbers are vendor-run and date from 2021: Linkerd 2.10.2 vs Istio 1.10.0 sidecars. Proxy memory was about 8x lower and median added latency was about half. The numbers and conditions are in Reference: Published Benchmark Numbers. No published benchmark compares Linkerd with Istio ambient mode, which removes per-pod sidecars. Resource comparisons against ambient mode therefore depend on the workload, and you should measure them yourself.
Scaling pressure falls mainly on the destination controller (memory grows with endpoints and churn) and on Prometheus (metric cardinality). Linkerd 2.20's destination refactor and the opt-in hostname labels on metrics both address these.
Linkerd vs Istio: Security and Design¶
| Aspect | Linkerd 2.20 | Istio (sidecar / ambient) |
|---|---|---|
| mTLS default | On between meshed pods. Plaintext from unmeshed clients is accepted unless policy denies it. | PERMISSIVE by default. Use PeerAuthentication STRICT to require mTLS. |
| Key exchange | Hybrid ML-KEM-768 + X25519 by default (since 2.19) | Depends on version and configuration. See the Istio notes. |
| Workload identity | ServiceAccount-derived DNS-style name. SPIFFE/SPIRE for non-Kubernetes workloads. | SPIFFE URI (spiffe://<td>/ns/<ns>/sa/<sa>) |
| Authorization CRDs | Server, HTTPRoute/GRPCRoute, AuthorizationPolicy, MeshTLSAuthentication, NetworkAuthentication | AuthorizationPolicy |
| End-user JWT validation | Not built in | RequestAuthentication |
| Sidecar-less mode | None. The project argues per-pod proxies are a stronger security boundary. | Ambient (ztunnel + waypoint) |
| Extensibility | None (fixed feature set) | Envoy filters, Wasm |
| Egress control | In-sidecar (EgressNetwork + Gateway API routes) |
Egress gateways, ServiceEntry, Sidecar resources |
| Rate limiting | Local, inbound (HTTPLocalRateLimitPolicy) |
Local and global (Envoy rate limit service) |
Key insight
Linkerd trades the extensibility and breadth of an Envoy-based mesh for a fixed, opinionated feature set. It needs almost no configuration to deliver mTLS, golden metrics and latency-aware balancing. Since 2024 it has closed several gaps it used to have: egress, rate limiting, circuit breaking with rate-limit awareness, federated multicluster, and native sidecars. The remaining gaps are sidecar-less operation, end-user JWT authentication, custom proxy extensions, and open-source stable release artifacts.
Sources¶
- Linkerd Architecture
- Automatic mTLS
- IPTables Reference
- Authorization Policy
- Circuit Breaking
- Load Balancing
- Multi-cluster communication
- Egress
- Native sidecars
- Announcing Linkerd 2.15, 2.16, 2.17, 2.18, 2.19, 2.20
- Why Linkerd doesn't use Envoy
- linkerd2-proxy repository
- CNCF: Linkerd graduation announcement
- The New Stack: Buoyant revises release model