Skip to content

OpenTelemetry

Summary

OpenTelemetry (OTel) is the CNCF's vendor-neutral standard for generating, collecting, and exporting telemetry: APIs, SDKs, the OTLP protocol, semantic conventions, the Collector, and the Kubernetes Operator. It became a CNCF Graduated project in May 2026. Traces, metrics, and logs are stable. Profiles are in public alpha. This topic distills the official guidance program into reusable architecture patterns: the "Managed Telemetry Platforms for Kubernetes Workloads" blueprint, the newer non-Kubernetes blueprint, and the three production reference implementations (Adobe, Mastodon, Skyscanner). The Collector engine itself is covered in the sibling OTel Collector topic.

Key Facts

Fact Value
Governance CNCF Graduated (moved to Graduated 2026-05-11, announced 2026-05-21)
License Apache-2.0
Latest Version (spec) v1.61.0 (2026-09-14)
Semantic conventions v1.44.0 (2026-08-04)
OTLP protobuf 1.11.0 (2026-07-21)
Operator v0.159.0 (2026-09-14). Kubernetes 1.25-1.36
Collector v0.161.0. See OTel Collector
Signal status Traces, metrics, logs, baggage: Stable. Profiles: Alpha (public alpha 2026-03-26)
Guidance program 2 published blueprints (1 more in progress), 3 reference implementations, run by the End-User SIG

Full status tables are in Reference.

Overview

The OpenTelemetry docs group best-practice material into two categories under Guidance:

Category Nature Purpose
Blueprints Living documents, each scoped to one environment or problem Solve common OTel adoption challenges. You may need to combine several
Reference implementations Point-in-time snapshots of real organizations' pipelines Ground blueprints in real deployments rather than theory. Explicitly not actively maintained

The program was announced on 2026-05-12 as an End-User SIG initiative run with the Developer Experience (DevEx) SIG. DevEx drove the three reference implementations. The End-User SIG listed three blueprints as work in progress: centralized telemetry platform (sig-end-user #246, now published as the K8s-workloads blueprint), non-Kubernetes infrastructure (#245, now published), and Kubernetes observability (#247, not yet published as of 2026-09-25). The distinction matters in practice: blueprint guidance changes as the project evolves, while a reference implementation records how one organization succeeded at one point in time.

This compact view shows the platform shape the managed-platforms blueprint recommends:

flowchart LR
    OP["OpenTelemetry Operator<br/>Instrumentation CR"] -.->|"injects SDK + defaults"| APP["Workload pods"]
    APP -->|"OTLP, no secrets"| GW["Collector Gateway tier<br/>k8s_attributes, filter,<br/>tail_sampling"]
    GW -->|"keys live here"| BE[("Backends")]

See Explanation for the full pattern breakdown, How-to Guides for deployment recipes, and Reference for versions, status matrices, and Operator annotations.

Reference Implementations

Organization Tagline Signature pattern Scale
Adobe An OTel pipeline designed for simplicity at scale Three-tier pipeline with immutable sidecar + signal-isolated managed namespace Thousands of collectors per signal, voluntary adoption
Mastodon Running Collectors in production with a small team One all-signals Collector per namespace as an Operator CR, deployed by Argo CD ~300k daily actives, ~10M req/min
Skyscanner Managing Collectors across 24 production clusters Single DNS entry point + Istio nearest-available routing over Gateway/Agent tiers 1,000+ microservices, 24 clusters, 6 platform engineers

Evaluation

Quick pattern selection from the three cases plus the blueprints:

  • Platform-engineering org running heterogeneous multi-tenant clusters: follow the managed-platforms blueprint. Use centrally managed, horizontally scaled Collector Gateways, OTLP-only export, secrets only at the Gateway tier, and Operator-first auto-instrumentation.
  • VMs, bare metal, or containers without an orchestrator: follow the non-K8s blueprint. Use OpAMP-managed agents where mature enough, plus the same Gateway layer.
  • Want zero restart risk for instrumented applications: Adobe's immutable sidecar config is designed so config changes never restart application pods.
  • Tiny team, global scale, minimal moving parts: Mastodon shows that one Collector per namespace behind the Operator and GitOps is "more than sufficient", with no tiers and no custom tooling.
  • Migrating off fragmented vendors to a single backend without lock-in: Skyscanner's vendor-agnostic DNS endpoint keeps collector topology stable across backend swaps. They moved 300+ services off OpenTracing by bumping one library version.

Adoption caveats

  • Profiles are alpha. The Profiling SIG advises against critical production use.
  • Messaging and GenAI semantic conventions are still Development. GenAI moved to its own repository in semconv v1.42.0, so expect renames.
  • The Collector ships every two weeks with no LTS, and the Operator can rewrite Collector CRs on upgrade. Budget for regular upgrades (see Upgrade Gotchas).

Topic Map

  • How-to Guides: build a managed OpenTelemetry platform on Kubernetes, following the blueprints and reference implementations.
  • Reference: signal and protocol stability, semantic-convention stability, Operator facts, catalogue of blueprints and reference implementations.
  • Explanation: how the parts fit together, why the guidance program recommends its topologies, how Adobe, Mastodon and Skyscanner adapted them.

The patterns here produce telemetry that every backend topic in this domain consumes:

  • SigNoz: OpenTelemetry-native unified platform on ClickHouse
  • LGTM Stack: Grafana's Tempo/Mimir/Loki receiving OTLP directly
  • Victoria Stack: VictoriaTraces/Logs accepting OTLP as drop-in replacements
  • Apache SkyWalking: accepts OTLP alongside its own agent format
  • Coroot: eBPF-first platform that also accepts OTLP spans
  • Observability 2.0: the wide-event critique of exactly the three-pillar pipeline this guidance standardizes

Related topics in other domains:

  • Kubernetes: the platform the Operator and Gateway patterns run on
  • Istio: Skyscanner's nearest-collector routing and span source, and one of the blueprint's gRPC load-balancing options
  • Argo CD and Flux CD: GitOps delivery of Collector CRs (Mastodon, Skyscanner, blueprint Action 3)
  • AI Platform Engineering: inference-service telemetry built on OTel conventions

Relation to the Domain Hub

The domain hub's Key Concepts section describes the Collector receiver, processor, exporter pipeline in general terms. This topic covers the deployment topologies (gateway tiers, sidecars, per-namespace CRs) that pipeline gets packaged into.

Sources

Primary pages were re-read from the open-telemetry/opentelemetry.io source on 2026-09-25. opentelemetry.io itself blocks this environment's fetcher, so links are the canonical published URLs.

Source Kind Note
Managed Telemetry Platforms for Kubernetes Workloads Primary, blueprint Core page researched here
Infrastructure and Processes in Non-K8s Environments Primary, blueprint Second published blueprint
Adobe reference implementation Primary, ref impl Byline 2026-04-08
Mastodon reference implementation Primary, ref impl Byline 2026-03-18
Skyscanner reference implementation Primary, ref impl Byline 2026-04-21
Guidance index Primary Defines blueprints vs reference implementations and the contribution process
Blueprints index Primary Lists published blueprints
Reference implementations index Primary Snapshot disclaimer
Introducing OTel Blueprints and Reference Implementations Primary, blog Program announcement, 2026-05-12
DevEx blog: Adobe · Mastodon · Skyscanner Primary, blog Companion posts
OpenTelemetry is a CNCF Graduated Project and CNCF announcement Primary Graduation, 2026-05-21
OpenTelemetry Profiles Enters Public Alpha Primary, blog 2026-03-26
Specification status and spec CHANGELOG Primary Signal stability, spec releases
opentelemetry-proto Primary OTLP maturity table and CHANGELOG
Semantic conventions CHANGELOG Primary Convention stability per area
Operator automatic instrumentation and Operator repo docs Primary, docs Annotation names, CRDs, compatibility matrix
Skyscanner Engineering on Medium Primary, company post Names New Relic and the 300+ service migration. Blocks non-browser clients, so open it in a browser
InfoQ: Skyscanner observability migration Secondary Independent corroboration

Questions

  • Which concrete backends does Adobe's routing-connector setup support today, and can teams switch backends purely through Helm values? The reference implementation says "multiple" without naming them.
  • How do these pipelines handle Gateway-tier failure modes such as persistent queuing and cross-regional failover when Istio's "nearest available" pool degrades? The blueprint's Appendix 4 covers file_storage trade-offs, but none of the three reference implementations reports running persistent queues.
  • When will the Kubernetes observability blueprint (sig-end-user #247) publish, and will the non-K8s blueprint gain its first reference architecture?
  • When will profiles move from alpha to beta, and which backends will support OTLP profiles in production?