Skip to content

Architecture

How OpenTelemetry's parts fit together, why the official guidance program recommends the topologies it does, and how three organizations adapted those ideas in production. The primary sources are the two published blueprints and three reference implementations on opentelemetry.io, re-read on 2026-09-25 from the site's source repository. Version and status facts sit in Reference. Runnable recipes sit in How-to Guides. The full link table is in index.

How the Pieces Fit

OpenTelemetry is a specification plus implementations. Applications call the API (a no-op until an SDK is registered). The SDK batches, samples, and exports over OTLP. The Collector receives, processes, and routes telemetry to backends. Semantic conventions give every signal the same attribute names, which is what makes correlation across signals possible. On Kubernetes, the Operator automates the Collector and SDK layers.

This diagram shows the components the guidance program assumes, with ownership split between application and platform teams:

flowchart LR
    subgraph AppOwned["Application team"]
        CODE["Business code"] --> API["OTel API"]
        LIBS["Instrumentation libraries<br/>or zero-code agent"] --> API
        API -.->|"implemented by"| SDK["OTel SDK<br/>BatchSpanProcessor<br/>PeriodicMetricReader"]
    end
    subgraph PlatformOwned["Platform team"]
        OP["OpenTelemetry Operator"]
        INST["Instrumentation CR<br/>base SDK config"]
        GW["Collector Gateway tier<br/>memory_limiter, k8s_attributes,<br/>filter, tail_sampling"]
    end
    SEMCONV[("Semantic conventions<br/>service.name, k8s.*, http.*")]
    OP -->|"webhook injects agent + env"| LIBS
    INST -.-> OP
    SDK -->|"OTLP 4317 gRPC / 4318 HTTP"| GW
    GW -->|"OTLP or vendor exporter"| BE[("Observability backends")]
    SEMCONV -.-> SDK
    SEMCONV -.-> GW

Traces, metrics, and logs are stable in the specification and in OTLP. Profiles entered public alpha in March 2026 and are not yet recommended for critical production use (see Reference).

Blueprints Versus Reference Implementations

The Guidance section holds two kinds of documents:

  • Blueprints are living documents. Each one is scoped to one environment, lists the challenges it solves, gives general guidelines, then gives implementation actions that link to the existing docs. Blueprints can overlap, extend, or relate to each other. They do not repeat how to configure an SDK or deploy a Collector, since the docs already cover that.
  • Reference implementations are point-in-time snapshots of how one organization did it. They are explicitly "not actively maintained" and may drift from current best practice.

The announcement (2026-05-12) frames the problem using Fred Brooks's terms. OTel adoption has essential complexity (breadth across the stack, backend neutrality) and accidental complexity (teams adopting it without a shared strategy, for example SDK configs incompatible with another team's Gateway, or mismatched propagators). Blueprints target the accidental part. The End-User SIG runs the program. The Developer Experience SIG drove the first three reference implementations.

This diagram shows how the published pages relate:

flowchart BT
    RA["Adobe RI<br/>2026-04-08"] -->|implements| BK["Blueprint: Managed Telemetry<br/>Platforms for K8s Workloads"]
    RM["Mastodon RI<br/>2026-03-18"] -->|implements| BK
    RS["Skyscanner RI<br/>2026-04-21"] -->|implements| BK
    BN["Blueprint: Infrastructure and<br/>Processes in Non-K8s Environments"] -.->|"relates to (shared Gateway layer)"| BK
    BO["Kubernetes observability<br/>blueprint - in progress"] -.->|"assumes central Gateway"| BK

Managed Telemetry Platforms for Kubernetes Workloads

The blueprint targets platform-engineering organizations running Kubernetes with highly autonomous product teams ("you build it, you run it"). It treats observability as an internal platform product: a paved road that gives every workload good default telemetry while teams keep ownership of the domain-specific parts.

The target architecture puts SDK defaults in the platform's hands and pushes all export through a central Gateway tier:

flowchart LR
    subgraph Workloads["Application workloads"]
        P1["Pod A<br/>Operator-injected auto-instrumentation"]
        P2["Pod B<br/>zero-code base image"]
        P3["Pod C<br/>shared language library"]
    end
    subgraph GatewayTier["Central Collector Gateway tier"]
        G1["Gateway 1<br/>horizontally scaled"]
        G2["Gateway 2"]
        S[("Backend endpoints<br/>and API keys<br/>live ONLY here")]
    end
    B1["Observability backend"]
    B2["SaaS backend"]
    P1 -->|"OTLP"| G1
    P2 -->|"OTLP gRPC"| G2
    P3 -->|"OTLP HTTP/protobuf"| G1
    G1 --> B1
    G2 --> B2
    S -.->|config| G1
    S -.->|config| G2

The five core prescriptions (quotes checked against the page source on 2026-09-25):

  1. Automatic centralized ingestion. Guideline 3: "We recommend that telemetry in this type of Kubernetes environment is automatically ingested into a centralized layer deployed as an OpenTelemetry Collector Gateway." Gateways can be chained, for example local per-cluster Gateways feeding a global tail-sampling Gateway, or namespace Gateways feeding a cluster Gateway.
  2. OTLP as the export protocol. Action 2's minimum base config: OTLP HTTP/protobuf (default) or OTLP gRPC "configured to export to the most optimal Collector (e.g. local Gateway in the same cluster)".
  3. Secrets never at the application layer. Action 2 note: "Backend/SaaS endpoints or API keys should not be included in application-level configuration, as we recommend handling these at a Collector Gateway." Challenge 3 describes direct egress from hundreds of applications as a single point of failure that removes central network governance.
  4. Operator-first instrumentation. Action 1: install the Operator, create Instrumentation CRs, annotate pods or namespaces. When the Operator is not possible, use zero-code base container images (supported languages) or shared language-specific libraries (other languages), and standardize on declarative configuration where the language supports it.
  5. Deploy Gateways with the Operator or the Helm charts (Action 3), stored as code and rolled out with Argo CD or Flux.

The rest of the blueprint adds organization standards to the base config (W3C tracecontext propagation, resource detectors, a curated instrumentation set, and minimum resource attributes service.name, service.version, service.namespace or service.owner, and deployment.environment.name), Gateway processors in order (k8s_attributes, then filter/transform/redaction, then tail_sampling behind a first layer that routes by trace ID with the load_balancing exporter), and self-monitoring of SDKs and Collectors (Action 5).

Compute-Efficiency Argument

The blueprint argues centrally managed Gateways use compute more efficiently than per-node DaemonSets or per-pod Sidecars in heterogeneous multi-tenant environments: a DaemonSet must be over-provisioned for variable node sizes ("a single node may serve 4 or 40 application pods") and fluctuating per-pod volume, while a Gateway tier scales independently against total telemetry volume. This is an efficiency argument inside the blueprint's scope, not a ban on agent or sidecar patterns. Skyscanner's DaemonSet below is a working counterexample doing deliberately narrow work.

Declared out of scope by the blueprint itself: audit logging and business reporting pipelines, compliance and inter-layer encryption, semantic-convention governance (it points to Weaver), and a full sampling architecture.

Gateway Design Trade-offs

The blueprint's appendices explain several choices that are easy to get wrong:

Decision Option A Option B Blueprint's view
Per-signal reliability Isolated Gateway per signal Several memory_limiter processors with different thresholds on one Gateway Isolation simplifies capacity planning but duplicates processor config. Multiple limiters rely on the OTLP receiver returning retryable errors so low-priority pipelines apply backpressure first
gRPC load balancing Client-side dns:/// plus round_robin against a headless Service L7 proxy or service mesh Kubernetes Services balance connections, so one long-lived HTTP/2 connection pins all traffic to one pod. Use the mesh if present, client-side balancing in-cluster, max_connection_age recycling as the simplest fallback, or OTLP/HTTP
Durability In-memory sending_queue file_storage-backed persistent queue Persistent queues need a StatefulSet with PVCs, and they delay backpressure because data goes to disk before memory pressure builds
Kubernetes enrichment k8s_attributes in the Gateway Downward API attributes set in the SDK In a Gateway every replica caches metadata for the whole cluster, and proxies hide the pod IP unless pass-through is on. Missing RBAC get/watch/list silently skips enrichment
Autoscaling signal CPU Memory, queue depth, active connections Use internal Collector telemetry, not default CPU-based HPA

Operator Injection Flow

Guideline 1 depends on the Operator's mutating webhook. The webhook changes pods only when they are created, which is why annotating a running Deployment has no effect until its pods restart:

sequenceDiagram
    participant Dev as Deploy pipeline
    participant API as kube-apiserver
    participant WH as Operator webhook
    participant Pod as New pod
    participant GW as Collector Gateway
    Dev->>API: Create pod with inject-java annotation
    API->>WH: Admission review
    WH->>WH: Resolve Instrumentation CR from annotation value
    WH-->>API: Patch pod - init container, shared volume, OTEL_* env
    API->>Pod: Schedule patched pod
    Pod->>Pod: Init container copies agent from autoinstrumentation image
    Pod->>Pod: App starts with agent and base SDK config
    Pod->>GW: OTLP export to endpoint from the CR

The sidecar annotation works the same way: the webhook adds a Collector container built from an OpenTelemetryCollector in mode: sidecar. Changing that sidecar's config means restarting application pods, which is the problem Adobe designed around.

Infrastructure and Processes in Non-K8s Environments

The second published blueprint covers VMs, bare metal, on-premises hosts, and containers run without an orchestrator. It names three challenges: limited automation for telemetry deployment, fragmented instrumentation approaches, and siloed data processing and export.

Its three guidelines mirror the Kubernetes blueprint without the Operator:

  1. Centrally manage agent lifecycle, allowing controlled customization. Use OpAMP where supported. The blueprint notes the OpAMP specification is Beta, so implementations should be evaluated first. Otherwise use shared libraries, pre-baked images, and existing config-management tooling.
  2. Centralize collection and processing through a Collector gateway layer.
  3. Standardize resource attribution and distribute reusable instrumentation building blocks.

Implementation is a six-step roadmap: define a baseline telemetry standard with layered configuration, stand up an OpAMP management plane, package standardized agents and SDK bootstrap artifacts, deploy the gateway layer, enforce resource-attribution standards, and centralize governance. Its reference-architecture list reads "Coming soon!" on 2026-09-25.

Adobe

Adobe runs its pipeline in three tiers:

flowchart TB
    subgraph ServiceNS["Service team namespace"]
        subgraph Pod["Application pod"]
            APP["App container"] --- SC["Sidecar Collector<br/>IMMUTABLE config"]
        end
        DC["Deployment Collector<br/>per service"]
    end
    subgraph ManagedNS["Managed namespace - observability team"]
        AUTH["Custom auth circuit-breaker<br/>extension on receiver"]
        M_METRICS["Collector Deployment<br/>metrics only"]
        M_LOGS["Collector Deployment<br/>logs only"]
        M_TRACES["Collector Deployment<br/>traces only"]
        RC["routing connector<br/>keyed on OTLP HTTP header"]
    end
    BE1["Backend chosen by team A"]
    BE2["Backend chosen by team B"]

    SC -->|"OTLP"| DC
    DC --> M_METRICS & M_LOGS & M_TRACES
    AUTH -.-> M_METRICS
    M_METRICS --> RC
    M_LOGS --> RC
    M_TRACES --> RC
    RC --> BE1
    RC --> BE2

All three tiers follow one rule: configuration changes must never restart application pods.

  • Tier 1: a user-facing Helm chart creates two collectors per service. The sidecar's config is locked down and "immutable to prevent application restarts caused by configuration changes". It collects every signal. A standalone Deployment collector receives sidecar telemetry over OTLP and is configurable through Helm values. "When configuration changes, only the deployment collector restarts. The application pod and its sidecar remain untouched."
  • Tier 2: the observability-team namespace runs a separate collector Deployment per signal (metrics, logs, traces): "If a backend becomes rate-limited or starts rejecting data for one signal type, the others continue flowing uninterrupted." These Deployments have generally run at default replica counts despite thousands of upstream collectors.
  • Tier 3: the backends. Multiple backends are supported.

Backend selection is delegated to teams. Helm values set an HTTP header on OTLP exports, and the managed-namespace collectors pass it to the routing connector. Adobe first used the routing processor and moved to the connector when the processor was deprecated. See How-to Guides.

Onboarding is two annotations, instrumentation.opentelemetry.io/inject-java and sidecar.opentelemetry.io/inject, with the Operator in every cluster. "People add two lines in their deployment. And it just works." Adoption is voluntary. Existing applications with established monitoring were not migrated.

Chained-collector error visibility. An OTLP hop returns 200 before the next tier tries the backend, so users "would just see 200s" while the backend rejected data. Adobe built a custom extension on the managed-namespace receiver. It sends mock authentication requests to the backend, caches the results, and returns 401 upstream when auth fails. It ships in Adobe's own Collector distribution, which contains only the components they use (teams can switch to Contrib).

Upgrade friction. Adobe upgrades the distribution and Operator quarterly. When the Operator upgrades, it can rewrite OpenTelemetryCollector resources for new config expectations, which can stop teams' older collector versions from starting. The fix is upgrading the collector, but teams saw breakage without changing anything themselves.

Skyscanner

Skyscanner routes all telemetry through one DNS name and two collector tiers:

flowchart TB
    SVC["Services across 24 clusters<br/>1000+ microservices"] -->|"OTEL_EXPORTER_OTLP_ENDPOINT=otel.skyscanner.net"| DNS["Single DNS entry point"]
    DNS --> ISTIO["Istio nearest-available routing"]
    ISTIO --> GW["Gateway Collectors<br/>ReplicaSet - bulk processing<br/>traces + metrics"]
    ISTIO_2["Istio proxies"] -->|"Zipkin spans"| GW
    GW -->|"spanmetrics connector"| GW
    AG["Agent Collectors<br/>DaemonSet - minimal processing<br/>Prometheus scrape of node-exporter,<br/>kube-state-metrics, kubelet"]
    GW -->|"OTLP"| VENDOR["Commercial observability vendor"]
    AG -->|"OTLP"| VENDOR
  • Vendor-agnostic ingress. Skyscanner adopted OTel in 2021 while moving from an internal open-source stack to a commercial vendor, specifically to avoid lock-in. Services everywhere send to one central DNS endpoint (otel.skyscanner.net), and "Istio handles routing requests to the nearest available collector." The vendor is New Relic per Skyscanner's own engineering post. opentelemetry.io calls it only "the observability vendor".
  • Division of labor. Gateway ReplicaSets take bulk OTLP traffic (traces and metrics) and do most processing. Agent DaemonSets scrape Prometheus endpoints of open-source and platform services that do not speak OTLP.
  • Cardinality lesson. Istio's native metrics had "cardinality explosion issues that would overwhelm their Prometheus deployment". Istio now emits spans (Zipkin originally), which the Gateway transforms to semantic conventions and turns into HTTP metrics through the span-metrics connector. This gives platform-level metrics for services whose code Skyscanner does not own.
  • Spans yes, SDK metrics no. Because Istio-derived metrics already exist, the Java base image uses SDK views to drop http.* and rpc.* metric aggregations while keeping spans. Teams can re-enable specific metrics under renamed instruments to avoid double counting.
  • Opinionated agent. The shared Java base image disables all instrumentations (OTEL_INSTRUMENTATION_COMMON_DEFAULT_ENABLED=false) and enables a curated set. Python and Node.js use wrapper libraries.
  • Migration proof. 300+ microservices moved from OpenTracing to OTel "in a matter of weeks" by bumping one core-library version, with no instrumentation-code changes (Skyscanner's Medium post and InfoQ. Self-reported).
  • Operations. Contrib distribution (they now know Contrib is not recommended for production and plan custom builds). Upgrades about every six months, which bunches breaking changes. Progressive Argo CD promotion: Dev, then 3 Alpha, 8 Beta, and the remaining 13 production clusters. Some pipelines still use the older, unstable HTTP semantic conventions because updating the Istio-mapping transform rules is manual.

Mastodon

Mastodon runs one collector per namespace, managed by the Operator and deployed with Argo CD:

flowchart TB
    subgraph NS["Namespace mastodon-social"]
        APPS["Rails web + Sidekiq pods"] -->|"OTLP"| COL
        subgraph COL["Single Collector Deployment - Contrib"]
            ALL["traces/all pipeline<br/>resource, k8sattributes,<br/>resourcedetection, transform"]
            DDC["datadog/connector<br/>APM stats on 100% of spans"]
            SAMP["traces/sample pipeline<br/>tail_sampling: errors + 0.1%"]
            ALL --> DDC --> SAMP
        end
    end
    GIT["Git repo"] --> ARGO["Argo CD"] -->|"deploys/promotes"| CR["OpenTelemetryCollector CR v1beta1"]
    CR -.-> OP["OpenTelemetry Operator<br/>reconcile - restart - lifecycle"]
    OP -.-> COL
    SAMP --> DD["Datadog"]
    DDC -->|"metrics pipeline"| DD

This is the deliberately minimal counterpoint: one all-signals Collector per Kubernetes namespace, run as a plain Deployment.

  • Declared as opentelemetry.io/v1beta1 OpenTelemetryCollector custom resources. The Operator handles reconciliation, restarts, and lifecycle. "There are no separate gateway and agent tiers, no complex routing layers, and no custom deployment tooling." Argo CD provides GitOps promotion with Git-history auditability.
  • Scale: mastodon.social serves up to 300,000 daily active users and ~10M requests/min on 9-15 autoscaling nodes (16 cores, 64 GB each) running ~70-80 pods, and this design "has proven more than sufficient". mastodon.online uses the same setup. The whole organization is ~20 people, and one engineer runs observability.
  • Stats before sampling. The published config computes Datadog APM stats with the datadog/connector on all spans, then feeds the connector's output into a second traces pipeline that tail-samples. Rate and latency metrics stay accurate while only ~0.1% of successful traces ("a few dozen per minute") plus all error traces are kept. Metrics and logs are not sampled.
  • No strict CPU/memory limits are enforced on Collector pods. Tim Campbell: "if it ever does have any issue, it just restarts automatically."
  • Mastodon also exposes OTel as plain environment variables for every self-hosted Mastodon operator, who can export directly with the Ruby SDK, go through a Collector, or disable telemetry.

Cross-Case Comparison

Dimension Adobe Mastodon Skyscanner
Tiering Three tiers One collector per namespace Two tiers (Gateway RS + Agent DS)
Deploy vehicle User-facing Helm chart + managed namespace Operator CR + Argo CD GitOps Platform-managed behind single DNS/Istio, Argo CD promotion
Restart blast radius Only the Deployment collector restarts. Pods untouched Operator auto-restarts collector only Infrastructure decoupled from app deploys
Secrets location Managed namespace Per-namespace CR (Kubernetes Secret refs) Gateway tier
Onboarding abstraction Two annotations Apps export OTLP to namespace-local collector Base Docker image / wrapper library bump
Signal isolation Explicit, one Deployment per signal None needed at its scale Partial, scrape path separate
Distribution Custom build Contrib Contrib
Team model Voluntary self-service platform ~20-person org, one engineer on observability 6 engineers manage most of 24 clusters
Scale reported Thousands of collectors per signal 300k DAU, 10M req/min 1,000+ microservices

Unifying Themes

Across the blueprint and the three cases, the shared lessons are:

  1. Small platform teams can run centralized collection. The value comes from abstraction (annotations, base images, CRs), not headcount.
  2. OTLP is the transport between tiers everywhere. Skyscanner also ingests Zipkin from Istio and scrapes Prometheus at the edge.
  3. Signal isolation buys resilience where rate limits and backpressure are real risks (Adobe explicit, Skyscanner partial). The blueprint offers per-signal memory_limiter thresholds as a lighter alternative.
  4. The developer-facing surface stays tiny: two annotation lines, a library bump, or nothing.
  5. Restart isolation and immutability are design goals, achieved with locked sidecars, Operator reconciliation, or GitOps-only changes.
  6. Chained collectors hide failures. Adobe's 200-before-export problem and the blueprint's Action 5 both say to monitor internal Collector and SDK telemetry rather than trusting a successful OTLP hop.

Verification Notes

Epistemic Status

  • All operational figures are company-authored case-study numbers (self-reported, not audited): scale stats, Skyscanner's weeks-long migration, Mastodon's 0.1% sampling outcome.
  • Reference implementations are officially "snapshots in time" and "not actively maintained". This page reflects the page source as of 2026-09-25 (byline dates 2026-03-18 to 2026-04-21).
  • The managed-platforms blueprint is a living document and has grown since the first 2026-08-27 pass (it now has five challenges, four guidelines, five actions, and five appendices). Quotes above were re-checked against the current source.
  • Vendor naming asymmetry: New Relic is named only in Skyscanner's own posts. Mastodon's config names Datadog directly.