Skip to content

Explanation

Scope

How Envoy Gateway works and why it is built this way: the runner pipeline, the IR, deployment modes, policy attachment and merging, rate limiting, extension points, the security model, and where Agent Router (formerly Envoy AI Gateway) and the Ingress-NGINX retirement fit. Commands and YAML recipes are in How-to Guides. Version matrices, ports, and CRD tables are in Reference.

Overview

Envoy Gateway is an open-source project for managing Envoy Proxy as a standalone or Kubernetes-based application gateway. It translates Kubernetes Gateway API resources, plus its own policy CRDs, into Envoy xDS configuration and manages the lifecycle of a fleet of Envoy proxies. Unlike full service meshes (Istio, Linkerd), Envoy Gateway focuses on north-south traffic: ingress at the edge, and egress to external services through the Backend resource.

The project started in 2022 to give Envoy a simple, supported gateway control plane, instead of each vendor building its own on top of raw xDS. Its GOALS.md lists four objectives: an expressive API based on Gateway API, "batteries included" defaults, support for all environments (not only Kubernetes), and extensibility for vendors without competing with vendor products. v1.0 shipped on 2024-03-13. The v1.9 line (2026-08) is current.

Core Components

The envoy-gateway binary starts a set of runners that communicate through in-memory publish/subscribe maps ("watchables"). The diagram shows the runners started in internal/cmd/server.go and what each one talks to.

graph TB
    subgraph CP["envoy-gateway controller"]
        PRV["Provider runner<br/>(Kubernetes, File, or Custom)"]
        GAR["Gateway API runner<br/>(gatewayapi translator)"]
        XDR["xDS runner<br/>(xDS translator, cache, gRPC server)"]
        INF["Infra runner<br/>(Kubernetes, Host, or Remote)"]
        GRL["Global rate limit runner<br/>(optional)"]
        ADM["Admin, metrics, traces runners"]
    end

    subgraph DP["Data plane"]
        EP1["Envoy Proxy pod"]
        EP2["Envoy Proxy pod"]
        RLS["envoy-ratelimit<br/>Deployment"]
    end

    K8S["Kubernetes API server"]
    REDIS["Redis"]

    K8S -->|"watch GatewayClass, Gateway,<br/>xRoutes, EG policies"| PRV
    PRV -->|"ProviderResources"| GAR
    GAR -->|"xDS IR"| XDR
    GAR -->|"Infra IR"| INF
    GAR -->|"status updates"| PRV
    XDR -->|"delta xDS over mTLS :18000"| EP1
    XDR -->|"delta xDS over mTLS :18000"| EP2
    XDR -->|"xDS resources"| GRL
    GRL -->|"rate limit config over xDS"| RLS
    INF -->|"server-side apply Deployment,<br/>Service, ServiceAccount, HPA"| K8S
    EP1 -->|"ShouldRateLimit gRPC"| RLS
    RLS --> REDIS

Provider Runner

The provider runner fetches resources from the configured provider and publishes them. With the Kubernetes provider it runs controller-runtime watches on GatewayClass, Gateway, HTTPRoute, GRPCRoute, TLSRoute, TCPRoute, UDPRoute, ListenerSet, BackendTLSPolicy, ReferenceGrant, Services, EndpointSlices, Secrets, ConfigMaps, and the EG CRDs. It also writes status back: when the translators publish status, the provider runner updates the resources. It must start before the infra runner because it creates the Kubernetes client the infra runner uses.

Since v1.9 the watches on ListenerSet, GRPCRoute, TLSRoute, and BackendTLSPolicy are conditional. Clusters whose provider-managed Gateway API CRD bundle omits them (GKE's managed add-on, OpenShift's Ingress Operator set) no longer crash-loop.

Gateway API Runner

The Gateway API runner subscribes to provider resources and runs the gatewayapi translator. The translator resolves GatewayClass ownership, listener conflicts, route attachment, ReferenceGrants, and policy targets. It produces two outputs per Gateway (or per merged Gateway group): an xDS IR and an Infra IR. It also computes status conditions such as Accepted, Programmed, ResolvedRefs, and, since v1.9, the RouteRulesOverlap warning for routes silently shadowed by identical matches.

xDS Runner

The xDS runner subscribes to the xDS IR, translates it into Envoy protobuf resources (Listener, RouteConfiguration, Cluster, ClusterLoadAssignment, Secret), applies EnvoyPatchPolicies and extension-server hooks, and publishes a snapshot into the xDS cache. The same runner hosts the gRPC xDS server on port 18000. Envoy bootstraps with api_type: DELTA_GRPC, so updates are incremental (delta ADS). The server authenticates proxies over mTLS. v1.9 fixed an authentication bypass in GatewayNamespace mode and raised the default max receive message size from 4 MiB to 32 MiB for large reconnect requests.

Infra Runner

The infra runner subscribes to the Infra IR and reconciles the proxy infrastructure. With the Kubernetes infrastructure provider it server-side-applies a Deployment (or DaemonSet), Service, ServiceAccount, optional HPA and PodDisruptionBudget, and the envoy-ratelimit Deployment when global rate limiting is on. With the Host provider (standalone mode) it runs Envoy as a local process. v1.9 added a Remote infrastructure provider so operators can plug in their own infrastructure management strategy.

Global Rate Limit Runner

When rateLimit is set in the EnvoyGateway config, this runner translates rate limit descriptors from the xDS resources into configuration for the envoyproxy/ratelimit service and serves it over xDS. The rate limit service keeps shared counters in Redis. Since v1.9.1 Envoy reaches the rate limit service through an EDS cluster built from the envoy-ratelimit EndpointSlices instead of a static DNS name, so scaling the service is picked up correctly.

Watchable metrics

All runners use the watchable pub/sub pattern. The controller exports watchable_depth, watchable_subscribe_duration_seconds, and watchable_publish_total to monitor internal event processing. v1.9.1 changed the bucket boundaries of watchable_subscribe_duration_seconds, so dashboards that use _bucket series need updating.

xDS Configuration Flow

This sequence shows what happens between kubectl apply and Envoy serving the new route.

sequenceDiagram
    participant Op as Cluster operator
    participant K8s as Kubernetes API
    participant PR as Provider runner
    participant GA as Gateway API runner
    participant XR as xDS runner
    participant IR as Infra runner
    participant EP as Envoy Proxy

    Op->>K8s: apply Gateway, HTTPRoute, SecurityPolicy
    K8s-->>PR: watch events
    PR->>GA: publish ProviderResources
    GA->>GA: resolve attachment, ReferenceGrants, policy targets
    GA->>IR: publish Infra IR
    IR->>K8s: server-side apply Envoy Deployment and Service
    GA->>XR: publish xDS IR
    XR->>XR: translate IR to LDS, RDS, CDS, EDS, SDS
    XR->>XR: apply EnvoyPatchPolicy and extension hooks
    XR->>EP: delta xDS push over gRPC
    EP-->>XR: ACK, or NACK with error detail
    GA->>PR: publish route and policy status
    PR->>K8s: update status conditions

A NACK means Envoy rejected the update and kept its last good config. v1.9 exports the xdsNACKTotal metric (labeled by node ID and type URL) so rejected configs no longer go unnoticed.

Translation Pipeline and IR

Envoy Gateway uses an Intermediate Representation to decouple Kubernetes-specific resource parsing from Envoy xDS generation. The pipeline has two branches.

flowchart LR
    RES["Gateway API and EG resources<br/>(Gateway, HTTPRoute, SecurityPolicy, ...)"] --> GAT["gatewayapi translator"]
    GAT --> IIR["Infra IR<br/>(proxy Deployment, Service,<br/>ports, EnvoyProxy settings)"]
    GAT --> XIR["xDS IR<br/>(listeners, routes,<br/>clusters, secrets)"]
    IIR --> INFM["Infra manager<br/>(Kubernetes, Host, Remote)"]
    XIR --> XDST["xDS translator"]
    XDST --> CACHE["xDS snapshot cache"]
    CACHE --> FLEET["Envoy Proxy fleet"]
    INFM --> FLEET
  • Infra IR describes the managed Envoy infrastructure: Deployment or DaemonSet, replicas, Service type and ports, and settings from the EnvoyProxy resource. The infra manager reconciles it.
  • xDS IR describes listeners, routes, clusters, and endpoints in a provider-agnostic form. The xDS translator converts it into Envoy protobuf resources.

This separation is what makes non-Kubernetes targets possible: standalone mode swaps the resource provider (File) and the infra provider (Host) while reusing both translators. The IR is internal and not a stable API, so extension servers and EnvoyPatchPolicies act on the final xDS instead.

The resource-to-xDS mapping table is in Reference.

Deployment Modes

Mode Resource provider Infra provider Status
Kubernetes (default) Kubernetes API Kubernetes, proxies in envoy-gateway-system Production
GatewayNamespace mode Kubernetes API Kubernetes, proxies in each Gateway's own namespace Supported; v1.9 fixed an xDS auth bypass and resource-ownership issues in this mode
Merged Gateways Kubernetes API One Envoy fleet shared by all Gateways of a GatewayClass (EnvoyProxy.spec.mergeGateways) Supported; listener ports must not conflict
Standalone File (watches a directory of YAML) Host (Envoy as a local process, or in a container) Experimental; upstream says not to use it in production
Remote infra (v1.9) Kubernetes API Remote, operator-defined infrastructure management New in v1.9

In standalone mode the controller runs as envoy-gateway server --config-path standalone.yaml. envoy-gateway certgen --local creates the TLS material that the runners use to talk to each other. Any change under the watched paths triggers an update. The Backend API must be enabled so routes can point at local IPs or Unix sockets. Directory layout follows XDG defaults (~/.config/envoy-gateway, ~/.local/share/envoy-gateway). The how-to guide has the full standalone recipe.

EnvoyProxy Configuration Hierarchy

The EnvoyProxy resource customizes the data plane. It is resolved at three levels (highest priority first):

  1. Gateway-level EnvoyProxy, referenced through Gateway.spec.infrastructure.parametersRef
  2. GatewayClass-level EnvoyProxy, referenced through GatewayClass.spec.parametersRef
  3. Default EnvoyProxy spec in the EnvoyGateway configuration (envoyProxy defaults, added in v1.8)

By default the most specific configuration replaces the others (mergeType: Replace). Since v1.8 an EnvoyProxy can set mergeType to StrategicMerge or JSONMerge so a Gateway-level resource layers on top of the class or controller defaults. Helm-configured proxy image settings only merge with an EnvoyProxy that uses a merge type.

Envoy Proxy Fleet

Each managed Envoy proxy runs as a Kubernetes Deployment (or DaemonSet) created by the infra runner:

  • Replicas: default 1 (DefaultDeploymentReplicas). Scale with EnvoyProxy.spec.provider.kubernetes.envoyDeployment.replicas or an HPA. Since v1.9 the controller no longer overrides HPA-computed replicas.
  • Bootstrap: generated by EG and passed to the pod. Dynamic configuration comes over delta xDS from port 18000.
  • Health: readiness is /ready on port 19003. Prometheus stats are at /stats/prometheus on port 19001. The admin interface listens on 127.0.0.1:19000 only.
  • Shutdown: a shutdown-manager sidecar coordinates graceful drain (shutdown.drainTimeout, 60s by default). v1.9 added healthCheckFailureDelay so draining can start before health checks fail.
  • Image: each EG minor pins one Envoy minor, for example distroless-v1.39.x for v1.9. Overriding image with another minor is untested.

Policy Attachment Model

Policies attach to Gateway API resources through targetRefs (by name) or targetSelectors (by label). This flowchart shows which policy kinds attach where.

flowchart TB
    GC["GatewayClass"] --> GW["Gateway<br/>(listeners)"]
    LS["ListenerSet"] -->|"parentRef"| GW
    GW --> HR["HTTPRoute / GRPCRoute"]
    GW --> TR["TCPRoute / UDPRoute / TLSRoute"]
    HR --> SVC["Service or Backend"]

    EPX["EnvoyProxy"] -.->|"parametersRef"| GC
    EPX -.->|"infrastructure.parametersRef"| GW
    EPP["EnvoyPatchPolicy"] -.->|"targetRef"| GW
    CTP["ClientTrafficPolicy"] -.->|"targetRef"| GW
    CTP -.->|"targetRef (v1.9)"| LS
    SP["SecurityPolicy"] -.->|"targetRef"| GW
    SP -.->|"targetRef"| HR
    BTP["BackendTrafficPolicy"] -.->|"targetRef"| GW
    BTP -.->|"targetRef"| HR
    BTP -.->|"targetRef"| TR
    EEP["EnvoyExtensionPolicy"] -.->|"targetRef"| HR
    BTLS["BackendTLSPolicy"] -.->|"targetRef"| SVC

Rules that follow from this model:

  • Specificity wins. A policy on a Gateway applies to every route on it. A policy of the same kind on a route or listener (sectionName) overrides the Gateway-level one for that traffic.
  • Merging is opt-in. Since v1.8 (SecurityPolicy) and earlier (BackendTrafficPolicy), a route-level policy can set mergeType so it merges with the parent policy instead of replacing it. v1.9 added mergeType to EnvoyExtensionPolicy. v1.9 also restricted mergeType to xRoute targets; it is rejected on Gateways and ListenerSets.
  • Cross-namespace attachment (v1.8) lets a policy target resources in another namespace when a ReferenceGrant allows it.
  • Status is written to the policy. Each policy reports conditions such as Accepted per ancestor, including when a more specific policy overrides it. v1.9 capped status.ancestors during translation after quadratic slowdowns with many policies on one target.

Rate Limiting Architecture

Rate limits are configured on BackendTrafficPolicy.spec.rateLimit. The older RateLimitFilter CRD from pre-1.0 releases no longer exists.

Mode Scope Backend Configuration
Local Per Envoy replica In-memory token bucket rateLimit.local.rules
Global Across all replicas envoy-ratelimit Deployment plus Redis rateLimit.global.rules, and rateLimit.backend in the EnvoyGateway config

Local limits multiply with the replica count: 100 req/min per replica across 3 replicas allows about 300 req/min in total. Global limits give one shared counter but add a gRPC round trip to the rate limit service on each matching request, plus Redis as a dependency. Both can be combined on one policy, and both support shadow mode (log-only). Field details are in Reference.

Extension Points

Envoy Gateway offers several extension layers, ordered here from safest to most powerful:

  • HTTPRouteFilter adds EG-specific route filters (regex path and host rewrites, direct responses, credential injection) without touching xDS.
  • EnvoyExtensionPolicy attaches data-plane extensions:
    • Wasm filters from HTTP or OCI sources. Since v1.9.1 OCI pulls require HTTPS unless a registry is marked insecure.
    • External Processing (ExtProc): a gRPC service that can read and mutate headers and bodies.
    • Lua: inline scripts. Disabled by default since v1.9 (extensionApis.enableLua) and sandboxed since v1.7.
    • Dynamic Modules (v1.8): shared-library extensions loaded into Envoy.
  • Backend CRD: FQDN, IP, and Unix-socket endpoints, plus DynamicResolver backends for forward-proxy use. It is disabled by default because it can route to arbitrary destinations.
  • Extension Manager: an external gRPC extension server that receives hooks during translation (route, virtual host, listener, translation-wide, and, since v1.8, PostEndpointsModify) and can rewrite the generated xDS. v1.8 allows several extension managers chained in sequence. v1.9 lets extension-server policies target HTTPRoutes, GRPCRoutes, and rules.
  • EnvoyPatchPolicy: JSON patches applied directly to generated xDS resources. It is the escape hatch for anything the API does not cover. It is disabled by default.

Extension servers and EnvoyPatchPolicy are full-trust

Upstream warns that enabling an extension server or EnvoyPatchPolicy may lead to complete compromise of the system. Whoever controls them can inject arbitrary configuration into every proxy. Generated xDS names also change between releases (for example JWT provider names in v1.9, Lua filters in v1.8.3), so patches and extension servers that match on names break across upgrades.

Security Model

Trust Boundaries

Actor Trusted with Main controls
Platform admin EnvoyGateway config, GatewayClass, EnvoyProxy, extension servers Cluster RBAC on envoy-gateway-system and cluster-scoped resources
Gateway owner Gateway listeners, TLS certificates, Gateway-level policies Namespace RBAC; allowedRoutes on listeners
Route owner HTTPRoute and route-level policies in their namespace ReferenceGrant for cross-namespace backends and policies
Client Nothing TLS, authentication, authorization, rate limits at the proxy

Most hardening in 2026 releases closes gaps between these roles. Lua was sandboxed (v1.7) and then disabled by default (v1.9). A tenant-supplied EnvoyProxy securityContext used to replace EG's hardened defaults, a confused-deputy issue (GHSA-w42f-28h3-998w, fixed in v1.9.1 and v1.8.4). Wasm OCI pulls could downgrade to plain HTTP (fixed in the same releases). The OIDC session cookie encryption moved from AES-256-CBC to AES-256-GCM to fix a padding oracle (CVE-2026-47775).

North-South Security Features

Envoy Gateway handles edge security. It does not manage east-west mTLS between in-cluster services.

This diagram shows where each control applies on a request's path.

flowchart LR
    C["Client"] -->|"TLS (Gateway listener<br/>certificateRefs)"| L["Envoy listener<br/>ClientTrafficPolicy"]
    L --> A["SecurityPolicy filters<br/>JWT, OIDC, API key,<br/>ExtAuth, CORS, CSRF"]
    A --> Z["Authorization rules<br/>(CIDR, GeoIP, CEL)"]
    Z --> R["Rate limit<br/>(BackendTrafficPolicy)"]
    R -->|"TLS or mTLS (BackendTLSPolicy,<br/>Backend tls)"| B["Backend"]
  • TLS termination. Certificates are Kubernetes Secrets referenced from Gateway listeners. mode: Terminate decrypts at the proxy. TLS passthrough uses a TLSRoute. Multiple certificates per listener are selected by SNI. Since v1.9, listener certificates can also come from an external SDS server (gateway.envoyproxy.io/sds Secrets, behind enableSDSSecretRef).
  • Client certificate validation is configured on ClientTrafficPolicy (tls.clientValidation), including CRLs. v1.9 added allowExpiredCertificate.
  • Backend TLS. BackendTLSPolicy (Gateway API v1 since v1.4) configures TLS origination and SAN validation towards a Service. The Backend CRD and EnvoyProxy.spec.backendTLS add a client certificate for backend mTLS. Since v1.9, backend TLS defaults to a maximum of TLS 1.3.
  • JWT validation supports remote and local JWKS, claim-to-header extraction, claim-based route recomputation (recomputeRoute), and, since v1.9, failOpen and failedRefetchDuration.
  • OIDC runs Envoy's oauth2 filter. v1.8 moved to a single filter per HCM with per-route config, and v1.9.1 rejects http:// issuers.
  • External authorization delegates the decision to a gRPC or HTTP service. The auth service receives request headers (and optionally the body), so it must be trusted and reached over TLS.
  • CORS and CSRF are SecurityPolicy fields. Gateway API's own CORS filter was not yet stable at the time of the ingress2gateway 1.0 release.

Setup recipes are in How-to Guides. The controls table is in Reference.

Request Flow

This sequence traces a single HTTPS request through the data plane, after configuration has been pushed.

sequenceDiagram
    participant Client as External client
    participant LB as Cloud load balancer
    participant Envoy as Envoy Proxy pod
    participant Auth as ExtAuth service
    participant RLS as envoy-ratelimit
    participant Backend as Backend Service

    Client->>LB: HTTPS request
    LB->>Envoy: forward (optionally with PROXY protocol)
    Envoy->>Envoy: TLS termination, client IP detection
    Envoy->>Envoy: route match on HTTPRoute rules
    Envoy->>Envoy: JWT or OIDC check (SecurityPolicy)
    Envoy->>Auth: ext_authz check (if configured)
    Auth-->>Envoy: allow
    Envoy->>RLS: ShouldRateLimit (global limits only)
    RLS-->>Envoy: OK
    Envoy->>Backend: forward to selected endpoint (LB policy, retries)
    Backend-->>Envoy: response
    Envoy-->>Client: response (response override, compression)

Performance Considerations

EG publishes a Nighthawk benchmark report with every release (route-scaling test on one CPU-limited proxy); the v1.9.1 figures are summarized in Reference. Envoy's per-request cost dominates the data path, so throughput and latency depend mostly on the filters enabled (ExtAuth, global rate limit, Wasm, and ExtProc each add work or a network hop) and on replica count. On the control plane, translation time and memory grow with the number of routes, policies, and EndpointSlices. v1.9 enabled EndpointSlice field indexing by default (faster lookups, more memory), added mergeBackends to deduplicate clusters, and fixed quadratic translation time when many policies target one resource.

Envoy AI Gateway and Agent Router

Envoy AI Gateway started in 2024 as an Envoy sub-project that adds LLM and MCP traffic handling on top of Envoy Gateway. On 2026-09-10 it joined the Agentic AI Foundation (AAIF) and was renamed Agent Router. The repository moved to theagentrouter/agent-router and the site to theagentrouter.ai. The aigateway.envoyproxy.io API group, the CRDs (AIGatewayRoute, AIServiceBackend, BackendSecurityPolicy), the aigw CLI, the envoy-ai-gateway-system namespace, container images, and Helm charts did not change. v1.0 shipped in June 2026 and v1.1 in August 2026.

Architecturally it is a second control plane that sits on Envoy Gateway. It writes Gateway API and EG resources (plus ExtProc configuration) that EG then turns into xDS. It recommends a two-tier pattern:

flowchart LR
    APP["Application<br/>(OpenAI-compatible client)"] --> T1["Tier One Gateway<br/>Agent Router on Envoy Gateway<br/>auth, routing, global rate limits"]
    T1 -->|"hosted models"| PROV["OpenAI, Azure OpenAI,<br/>AWS Bedrock, Vertex AI"]
    T1 -->|"self-hosted models"| T2["Tier Two Gateway<br/>Envoy Gateway plus<br/>Inference Extension endpoint picker"]
    T2 --> POD["Model server pods<br/>(vLLM and similar)"]
    T1 -->|"MCP"| MCP["MCP servers"]

Agent Router v1.0.x requires EG v1.8.1 or later, so AI gateway users are pulled onto the same quarterly EG upgrade cycle.

Ingress-NGINX Retirement Context

Kubernetes SIG Network announced the retirement of the community Ingress-NGINX controller on 2025-11-11. Best-effort maintenance ended in March 2026, after which there are no further releases or security fixes. The Steering and Security Response Committees issued a joint statement on 2026-01-29 urging users to migrate. Gateway API is the designated successor to Ingress.

Envoy Gateway is one of the main destinations for three reasons. It is vendor-neutral under the Envoy project. It covers many common Ingress-NGINX annotations (rewrites, CORS, auth, rate limits, timeouts) through typed policies instead of annotations. And ingress2gateway 1.0 (2026-03-20) ships an envoy-gateway emitter that outputs EG policy resources alongside core Gateway API objects. What does not translate is arbitrary NGINX configuration: configuration-snippet, server-snippet, and custom Lua need a redesign, usually with EnvoyExtensionPolicy or EnvoyPatchPolicy. Migration steps are in How-to Guides. See also Kubernetes how-to guides.

Design Trade-offs

  • Typed policies instead of annotations. Policies are validated by CEL and reported in status, which fixes the silent-failure mode of annotations. The cost is a large, still-v1alpha1 API surface with breaking changes in most minor releases.
  • One Envoy fleet per Gateway by default. This gives strong isolation between Gateways but more pods. mergeGateways trades isolation for density.
  • One Envoy minor per EG minor. This keeps generated config and proxy in lockstep. The cost is a 6-month support window that forces regular upgrades.
  • Dangerous features are off by default. EnvoyPatchPolicy, Backend, Lua, and external SDS secrets must be enabled explicitly in the EnvoyGateway config, which keeps the default install safe for multi-tenant clusters.

Key insight

Envoy Gateway separates resource watching, Gateway API translation, xDS translation and serving, and infrastructure management into distinct runners joined by an IR. That split is what lets the same translators serve Kubernetes, standalone hosts, and remote infrastructure, and what lets Agent Router build on EG instead of forking it.

Sources