OpenTelemetry¶
Summary
OpenTelemetry (OTel) is the CNCF's vendor-neutral standard for generating, collecting, and exporting telemetry: APIs, SDKs, the OTLP protocol, semantic conventions, the Collector, and the Kubernetes Operator. It became a CNCF Graduated project in May 2026. Traces, metrics, and logs are stable. Profiles are in public alpha. This topic distills the official guidance program into reusable architecture patterns: the "Managed Telemetry Platforms for Kubernetes Workloads" blueprint, the newer non-Kubernetes blueprint, and the three production reference implementations (Adobe, Mastodon, Skyscanner). The Collector engine itself is covered in the sibling OTel Collector topic.
Key Facts¶
| Fact | Value |
|---|---|
| Governance | CNCF Graduated (moved to Graduated 2026-05-11, announced 2026-05-21) |
| License | Apache-2.0 |
| Latest Version (spec) | v1.61.0 (2026-09-14) |
| Semantic conventions | v1.44.0 (2026-08-04) |
| OTLP protobuf | 1.11.0 (2026-07-21) |
| Operator | v0.159.0 (2026-09-14). Kubernetes 1.25-1.36 |
| Collector | v0.161.0. See OTel Collector |
| Signal status | Traces, metrics, logs, baggage: Stable. Profiles: Alpha (public alpha 2026-03-26) |
| Guidance program | 2 published blueprints (1 more in progress), 3 reference implementations, run by the End-User SIG |
Full status tables are in Reference.
Overview¶
The OpenTelemetry docs group best-practice material into two categories under Guidance:
| Category | Nature | Purpose |
|---|---|---|
| Blueprints | Living documents, each scoped to one environment or problem | Solve common OTel adoption challenges. You may need to combine several |
| Reference implementations | Point-in-time snapshots of real organizations' pipelines | Ground blueprints in real deployments rather than theory. Explicitly not actively maintained |
The program was announced on 2026-05-12 as an End-User SIG initiative run with the Developer Experience (DevEx) SIG. DevEx drove the three reference implementations. The End-User SIG listed three blueprints as work in progress: centralized telemetry platform (sig-end-user #246, now published as the K8s-workloads blueprint), non-Kubernetes infrastructure (#245, now published), and Kubernetes observability (#247, not yet published as of 2026-09-25). The distinction matters in practice: blueprint guidance changes as the project evolves, while a reference implementation records how one organization succeeded at one point in time.
This compact view shows the platform shape the managed-platforms blueprint recommends:
flowchart LR
OP["OpenTelemetry Operator<br/>Instrumentation CR"] -.->|"injects SDK + defaults"| APP["Workload pods"]
APP -->|"OTLP, no secrets"| GW["Collector Gateway tier<br/>k8s_attributes, filter,<br/>tail_sampling"]
GW -->|"keys live here"| BE[("Backends")]
See Explanation for the full pattern breakdown, How-to Guides for deployment recipes, and Reference for versions, status matrices, and Operator annotations.
Reference Implementations¶
| Organization | Tagline | Signature pattern | Scale |
|---|---|---|---|
| Adobe | An OTel pipeline designed for simplicity at scale | Three-tier pipeline with immutable sidecar + signal-isolated managed namespace | Thousands of collectors per signal, voluntary adoption |
| Mastodon | Running Collectors in production with a small team | One all-signals Collector per namespace as an Operator CR, deployed by Argo CD | ~300k daily actives, ~10M req/min |
| Skyscanner | Managing Collectors across 24 production clusters | Single DNS entry point + Istio nearest-available routing over Gateway/Agent tiers | 1,000+ microservices, 24 clusters, 6 platform engineers |
Evaluation¶
Quick pattern selection from the three cases plus the blueprints:
- Platform-engineering org running heterogeneous multi-tenant clusters: follow the managed-platforms blueprint. Use centrally managed, horizontally scaled Collector Gateways, OTLP-only export, secrets only at the Gateway tier, and Operator-first auto-instrumentation.
- VMs, bare metal, or containers without an orchestrator: follow the non-K8s blueprint. Use OpAMP-managed agents where mature enough, plus the same Gateway layer.
- Want zero restart risk for instrumented applications: Adobe's immutable sidecar config is designed so config changes never restart application pods.
- Tiny team, global scale, minimal moving parts: Mastodon shows that one Collector per namespace behind the Operator and GitOps is "more than sufficient", with no tiers and no custom tooling.
- Migrating off fragmented vendors to a single backend without lock-in: Skyscanner's vendor-agnostic DNS endpoint keeps collector topology stable across backend swaps. They moved 300+ services off OpenTracing by bumping one library version.
Adoption caveats
- Profiles are alpha. The Profiling SIG advises against critical production use.
- Messaging and GenAI semantic conventions are still Development. GenAI moved to its own repository in semconv v1.42.0, so expect renames.
- The Collector ships every two weeks with no LTS, and the Operator can rewrite Collector CRs on upgrade. Budget for regular upgrades (see Upgrade Gotchas).
Topic Map¶
- How-to Guides: build a managed OpenTelemetry platform on Kubernetes, following the blueprints and reference implementations.
- Reference: signal and protocol stability, semantic-convention stability, Operator facts, catalogue of blueprints and reference implementations.
- Explanation: how the parts fit together, why the guidance program recommends its topologies, how Adobe, Mastodon and Skyscanner adapted them.
Related Topics¶
- OTel Collector: sibling component topic covering the engine internals (pdata, distributions, autoscaling, tail-sampling placement) that these organizational patterns deploy.
- Domain comparisons: Observability Stacks Comparison (OTel as the collection layer across stacks) and LGTM vs Victoria Stack.
The patterns here produce telemetry that every backend topic in this domain consumes:
- SigNoz: OpenTelemetry-native unified platform on ClickHouse
- LGTM Stack: Grafana's Tempo/Mimir/Loki receiving OTLP directly
- Victoria Stack: VictoriaTraces/Logs accepting OTLP as drop-in replacements
- Apache SkyWalking: accepts OTLP alongside its own agent format
- Coroot: eBPF-first platform that also accepts OTLP spans
- Observability 2.0: the wide-event critique of exactly the three-pillar pipeline this guidance standardizes
Related topics in other domains:
- Kubernetes: the platform the Operator and Gateway patterns run on
- Istio: Skyscanner's nearest-collector routing and span source, and one of the blueprint's gRPC load-balancing options
- Argo CD and Flux CD: GitOps delivery of Collector CRs (Mastodon, Skyscanner, blueprint Action 3)
- AI Platform Engineering: inference-service telemetry built on OTel conventions
Relation to the Domain Hub
The domain hub's Key Concepts section describes the Collector receiver, processor, exporter pipeline in general terms. This topic covers the deployment topologies (gateway tiers, sidecars, per-namespace CRs) that pipeline gets packaged into.
Sources¶
Primary pages were re-read from the open-telemetry/opentelemetry.io source on 2026-09-25. opentelemetry.io itself blocks this environment's fetcher, so links are the canonical published URLs.
| Source | Kind | Note |
|---|---|---|
| Managed Telemetry Platforms for Kubernetes Workloads | Primary, blueprint | Core page researched here |
| Infrastructure and Processes in Non-K8s Environments | Primary, blueprint | Second published blueprint |
| Adobe reference implementation | Primary, ref impl | Byline 2026-04-08 |
| Mastodon reference implementation | Primary, ref impl | Byline 2026-03-18 |
| Skyscanner reference implementation | Primary, ref impl | Byline 2026-04-21 |
| Guidance index | Primary | Defines blueprints vs reference implementations and the contribution process |
| Blueprints index | Primary | Lists published blueprints |
| Reference implementations index | Primary | Snapshot disclaimer |
| Introducing OTel Blueprints and Reference Implementations | Primary, blog | Program announcement, 2026-05-12 |
| DevEx blog: Adobe · Mastodon · Skyscanner | Primary, blog | Companion posts |
| OpenTelemetry is a CNCF Graduated Project and CNCF announcement | Primary | Graduation, 2026-05-21 |
| OpenTelemetry Profiles Enters Public Alpha | Primary, blog | 2026-03-26 |
| Specification status and spec CHANGELOG | Primary | Signal stability, spec releases |
| opentelemetry-proto | Primary | OTLP maturity table and CHANGELOG |
| Semantic conventions CHANGELOG | Primary | Convention stability per area |
| Operator automatic instrumentation and Operator repo docs | Primary, docs | Annotation names, CRDs, compatibility matrix |
| Skyscanner Engineering on Medium | Primary, company post | Names New Relic and the 300+ service migration. Blocks non-browser clients, so open it in a browser |
| InfoQ: Skyscanner observability migration | Secondary | Independent corroboration |
Questions¶
- Which concrete backends does Adobe's routing-connector setup support today, and can teams switch backends purely through Helm values? The reference implementation says "multiple" without naming them.
- How do these pipelines handle Gateway-tier failure modes such as persistent queuing and cross-regional failover when Istio's "nearest available" pool degrades? The blueprint's Appendix 4 covers
file_storagetrade-offs, but none of the three reference implementations reports running persistent queues. - When will the Kubernetes observability blueprint (sig-end-user #247) publish, and will the non-K8s blueprint gain its first reference architecture?
- When will profiles move from alpha to beta, and which backends will support OTLP profiles in production?