Skip to content

How-to Guides

Task recipes for building a managed OpenTelemetry platform on Kubernetes, taken from the managed-platforms blueprint and the Adobe, Mastodon, and Skyscanner reference implementations. Fragments marked illustrative show documented component shapes, not byte-exact copies of any organization's internal config. Versions and annotation tables are in Reference. The reasoning behind each pattern is in Explanation.

Platform Setup Recipes

Install The Operator

The Operator's admission webhook needs TLS certificates. The default path uses cert-manager:

# cert-manager (the chart value crds.enabled installs its CRDs)
helm repo add jetstack https://charts.jetstack.io
helm install cert-manager jetstack/cert-manager \
  --namespace cert-manager --create-namespace --set crds.enabled=true

# Option 1: plain manifest from the latest Operator release
kubectl apply -f https://github.com/open-telemetry/opentelemetry-operator/releases/latest/download/opentelemetry-operator.yaml

# Option 2: Helm chart (the collector image repository must be set explicitly)
helm repo add open-telemetry https://open-telemetry.github.io/opentelemetry-helm-charts
helm install opentelemetry-operator open-telemetry/opentelemetry-operator \
  --namespace opentelemetry-operator-system --create-namespace \
  --set "manager.collectorImage.repository=ghcr.io/open-telemetry/opentelemetry-collector-releases/opentelemetry-collector-k8s"

Without cert-manager, the chart can generate a self-signed certificate: add --set admissionWebhooks.certManager.enabled=false --set admissionWebhooks.autoGenerateCert.enabled=true.

Check that the install worked:

kubectl get pods -n opentelemetry-operator-system
kubectl get crd | grep opentelemetry.io

Publish A Base Instrumentation CR

This is blueprint Action 1 and Action 2: one platform-owned Instrumentation CR carries the organization's SDK defaults. It sets W3C propagation, parent-based sampling, and an in-cluster Gateway endpoint, and it has no backend keys (illustrative names):

apiVersion: opentelemetry.io/v1alpha1
kind: Instrumentation
metadata:
  name: platform-default
  namespace: team-a
spec:
  exporter:
    endpoint: http://otel-gateway-collector.observability.svc:4318   # Gateway, never the SaaS URL
  propagators:
    - tracecontext
    - baggage
  sampler:
    type: parentbased_traceidratio
    argument: "1"
  resource:
    addK8sUIDAttributes: true
  env:
    - name: OTEL_RESOURCE_ATTRIBUTES
      value: deployment.environment.name=production,service.namespace=team-a

The Operator docs warn that the Python, .NET, and Go agents default to http/protobuf, so they need port 4318. Only point them at 4317 if you also switch them to gRPC.

kubectl apply -f instrumentation.yaml
kubectl get otelinst -A

Deploy A Gateway With The Operator

This is Action 3. The Gateway is an OpenTelemetryCollector in mode: deployment (or statefulset for persistent queues), with memory_limiter first in every pipeline (illustrative):

apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata:
  name: otel-gateway
  namespace: observability
spec:
  mode: deployment
  replicas: 3
  config:
    receivers:
      otlp:
        protocols:
          grpc: { endpoint: 0.0.0.0:4317 }
          http: { endpoint: 0.0.0.0:4318 }
    processors:
      memory_limiter:
        check_interval: 1s
        limit_percentage: 80
        spike_limit_percentage: 20
    exporters:
      otlp_http:
        endpoint: https://otlp.backend.example
        sending_queue:
          enabled: true
        retry_on_failure:
          enabled: true
    service:
      pipelines:
        traces:  { receivers: [otlp], processors: [memory_limiter], exporters: [otlp_http] }
        metrics: { receivers: [otlp], processors: [memory_limiter], exporters: [otlp_http] }
        logs:    { receivers: [otlp], processors: [memory_limiter], exporters: [otlp_http] }

Exporter names

Collector v0.144.0 renamed the otlphttp exporter to otlp_http and kept otlphttp as a deprecated alias. Use otlphttp on older Collectors. See the OTel Collector topic for version details.

Store this CR in Git and deploy it with Argo CD or Flux. That gives you an audit trail, phased rollouts, and one-step rollbacks.

Onboarding Recipes

Adobe-Style: Two Annotations

Prerequisites:

  1. The OpenTelemetry Operator is installed in every cluster (Adobe runs it cluster-wide).
  2. An Instrumentation CR exists for the target language or runtime.
  3. An OpenTelemetryCollector with mode: sidecar exists in the namespace. At Adobe, the team-facing Helm chart creates it.

Opt a Java workload in by adding exactly two annotations to the pod template:

# deployment.yaml - annotations go under spec.template.metadata, not the Deployment's metadata
spec:
  template:
    metadata:
      annotations:
        instrumentation.opentelemetry.io/inject-java: "true"
        sidecar.opentelemetry.io/inject: "true"

Wrong place, silent no-op

The Operator's sidecar docs call out annotations on the Deployment's own metadata as WRONG. The webhook only sees the pod. Injection also happens only when a pod is created, so run kubectl rollout restart deployment/<name> after annotating.

Use the matching segment for other runtimes (inject-nodejs, inject-python, inject-dotnet, inject-go, inject-apache-httpd, inject-nginx, inject-sdk). The value can be "true", a CR name, or namespace/name. At Adobe, telemetry flows through the locked-down sidecar to a Deployment collector whose config can change without touching the pod.

Mastodon-Style: One CR Per Namespace, GitOps Only

Declare one all-signals collector as an OpenTelemetryCollector custom resource, commit it, and let Argo CD deploy it while the Operator owns its lifecycle. This excerpt follows the published production config: secrets come from a Kubernetes Secret, APM stats are computed on every span, and sampling happens afterwards in a second traces pipeline. The otlp receiver and most processor definitions are left out, and processor lists are shortened.

apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata:
  name: mastodon-social
  namespace: mastodon-social
spec:
  env:
    - name: DD_API_KEY
      valueFrom: { secretKeyRef: { name: datadog-secret, key: api-key } }
  config:
    processors:
      tail_sampling:
        policies:
          [
            { name: errors-policy, type: status_code, status_code: { status_codes: [ERROR] } },
            { name: randomized-policy, type: probabilistic, probabilistic: { sampling_percentage: 0.1 } },
          ]
    connectors:
      datadog/connector:
        traces: { compute_stats_by_span_kind: true }
    exporters:
      datadog:
        api: { site: ${DD_SITE}, key: ${DD_API_KEY} }
    service:
      pipelines:
        traces/all:    { receivers: [otlp], processors: [resource, k8sattributes, batch], exporters: [datadog/connector] }
        traces/sample: { receivers: [datadog/connector], processors: [tail_sampling, batch], exporters: [datadog] }
        metrics:       { receivers: [datadog/connector, otlp], processors: [resource, k8sattributes, batch], exporters: [datadog] }
        logs:          { receivers: [otlp], processors: [resource, k8sattributes, batch], exporters: [datadog] }

Result at mastodon.social: every error trace is kept, and only "a few dozen" successful traces per minute survive (~0.1%). Metrics and logs are not sampled. Mastodon enforces no strict CPU/memory limits on the Collector: "if it ever does have any issue, it just restarts automatically." The full manifest is on the Mastodon reference implementation page.

Skyscanner-Style: Endpoint + Opinionated Base Image

The vendor contract lives in a shared Java base image. This excerpt comes from Skyscanner's published sample, which the page describes as an illustration:

FROM ghcr.io/open-telemetry/opentelemetry-operator/autoinstrumentation-java:2.25.0 AS otel
FROM image/registry/public-java-image:x.y.z
COPY --from=otel /javaagent.jar $OPEN_TELEMETRY_DIRECTORY/opentelemetry-javaagent.jar
ENV OTEL_EXPORTER_OTLP_ENDPOINT="http://otel.skyscanner.net"
ENV OTEL_LOGS_EXPORTER="none"
ENV OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE="DELTA"
ENV OTEL_EXPERIMENTAL_METRICS_VIEW_CONFIG="otel-view.yaml"
# Disable everything, then enable a curated set
ENV OTEL_INSTRUMENTATION_COMMON_DEFAULT_ENABLED="false"
ENV OTEL_INSTRUMENTATION_RUNTIME_TELEMETRY_ENABLED="true"
ENV OTEL_INSTRUMENTATION_APACHE_HTTPCLIENT_ENABLED="true"
  • Service teams inherit the base image. A service Dockerfile only adds extra OTEL_INSTRUMENTATION_<NAME>_ENABLED=true lines. Python and Node.js services use wrapper libraries instead.
  • A launcher script builds -Dotel.resource.attributes from deployment-provided variables (cloud.region, cloud.account.id, k8s.cluster.name, service.name).
  • Migrating legacy OpenTracing services meant bumping the core-library version, with no instrumentation-code changes.
  • Services that cannot emit OTLP are scraped by the Agent DaemonSet tier instead of being left out.

Configuration Recipes

Signal Isolation Per Backend Risk (Adobe Tier 2)

Run one collector Deployment per signal in the platform-managed namespace, so that a backend rate-limiting or rejecting one signal never blocks the others:

metrics-deployment ----\
logs-deployment --------+--> routing connector --> per-team exporter
traces-deployment -----/

The blueprint's lighter alternative is one Gateway with a separate memory_limiter per signal pipeline, where the lower-priority pipelines get lower thresholds so they apply backpressure first.

Team-Controlled Backend Routing Via Header (Adobe)

Service teams choose a backend through Helm values, and the value becomes an HTTP header on OTLP exports:

# values.yaml of the team-facing chart (illustrative names follow the described mechanism)
telemetry:
  backend: backend-a         # sets an HTTP header on OTLP exports

The managed-namespace collectors route on that header with the routing connector. The receiver must keep client metadata, and the old request["..."] context is deprecated in favor of otelcol.client.metadata (illustrative):

receivers:
  otlp:
    protocols:
      http: { endpoint: 0.0.0.0:4318, include_metadata: true }
connectors:
  routing:
    default_pipelines: [traces/default]
    table:
      - context: resource
        condition: otelcol.client.metadata["X-Backend"][0] == "backend-a"
        pipelines: [traces/backend-a]
service:
  pipelines:
    traces/in:        { receivers: [otlp], exporters: [routing] }
    traces/backend-a: { receivers: [routing], exporters: [otlp_http/backend-a] }
    traces/default:   { receivers: [routing], exporters: [otlp_http/default] }

Span Metrics Instead Of Native Istio Metrics (Skyscanner)

Skyscanner stopped relying on Istio's native metrics after cardinality explosions overwhelmed Prometheus. The Gateway ingests Istio spans (Zipkin receiver), maps them to semantic conventions, and derives lower-cardinality HTTP metrics with the span-metrics connector. The published settings include delta temporality, exponential histograms (max_size: 160, unit ms), dimensions_cache_size: 15000000, and a 30 s flush interval. The general rule: when series counts explode, derive metrics in the Collector rather than ingesting raw mesh metrics.

Drop Duplicate SDK Metrics With Views (Skyscanner)

To keep spans but drop SDK HTTP/RPC metrics that duplicate mesh-derived ones, point OTEL_EXPERIMENTAL_METRICS_VIEW_CONFIG (Java agent) at a view file:

- selector: { instrument_name: http.* }
  view: { aggregation: drop }
- selector: { instrument_name: rpc.* }
  view: { aggregation: drop }

A team that needs route-level latency adds a view that keeps http.server.request.duration under a new name (Skyscanner uses app.http.server.request.duration), so it cannot be double-counted against the Istio series.

Un-error Expected 404s Before Error-Based Sampling (Skyscanner)

Cache services that return 404 for "not found" caused 100% trace retention because the spans were errors. A span processor resets the status for those responses only (hostname regex shortened from the published one):

processors:
  span/unset_cache_client_404:
    include:
      match_type: regexp
      attributes:
        - { key: http.response.status_code, value: ^404$ }
        - { key: server.address, value: ^(service-x\.example\.net)$ }
    status:
      code: Unset

With SDK declarative configuration, the same filtering could now be owned by service teams instead of the central Collector.

Balance OTLP gRPC Across Gateway Replicas

A Kubernetes Service balances connections, not gRPC requests, so a long-lived HTTP/2 connection pins one sender to one Gateway pod. Pick one of these fixes (blueprint Appendix 2):

# 1. Client side: headless Service + DNS resolver (round_robin is the default since Collector v0.105.0)
exporters:
  otlp:
    endpoint: dns:///otel-gateway-collector-headless.observability.svc.cluster.local:4317
    balancer_name: round_robin
---
# 2. Server side: recycle connections so the Service redistributes them
receivers:
  otlp:
    protocols:
      grpc:
        keepalive:
          server_parameters:
            max_connection_age: 60s
            max_connection_age_grace: 10s

The Operator creates <name>-collector-headless automatically only in statefulset mode. For mode: deployment, create a Service with clusterIP: None yourself. If a mesh such as Istio or Linkerd is present, its sidecar already balances per request. OTLP/HTTP avoids the problem entirely.

Grant RBAC For k8s_attributes

The k8s_attributes processor silently skips attributes it cannot read. Give the Collector's ServiceAccount get, watch, and list on every resource it enriches from, then check:

kubectl auth can-i list deployments --as=system:serviceaccount:observability:otel-gateway-collector
kubectl auth can-i watch pods --as=system:serviceaccount:observability:otel-gateway-collector

In a Gateway, every replica caches metadata for the whole cluster, so memory grows with cluster size. If a proxy sits in front, enable pass-through so the Gateway sees the original pod IP, or set k8s.pod.uid from the Downward API and associate on that attribute.

Upgrade Gotchas

  • Routing processor to routing connector: if you adopted header-based backend routing early, migrate to the connector, because the processor was deprecated (Adobe took this path). Inside the connector, move from the deprecated request[...] context to otelcol.client.metadata[...].
  • Operator upgrades rewrite Collector CRs: a new Operator can modify OpenTelemetryCollector resources for new config expectations, which can stop teams' older collector images from starting (Adobe). Keep collector images close to the Operator's version, or let the Operator manage versions by not pinning spec.image.
  • Operator v0.158.0 turns on NetworkPolicies by default for the operator and operands. Check egress from Collectors and the Target Allocator after upgrading.
  • Operator v0.159.0 Target Allocator filter strategy: an empty filter_strategy now means none, and unknown values are rejected at startup.
  • Six-month upgrade gaps pile up breaking changes (Skyscanner). With two-week Collector releases and no LTS, smaller and more frequent upgrades are easier to review.
  • Semantic-convention renames: v1.44.0 renamed *.memory.paging.faults metrics to *.paging.faults, and v1.42.0 moved all gen_ai.* conventions to a separate repository. Pin the schema URL and review transform rules that hard-code names (Skyscanner still runs old HTTP conventions for this reason).
  • DaemonSet sizing on heterogeneous nodes: the blueprint's efficiency argument (nodes serving anywhere from 4 to 40 pods force over-provisioned DaemonSets) is a reason to re-evaluate agent tiers as node sizes diverge. It is not a rule to remove them (see Skyscanner's narrow-scrape DaemonSet).
  • Secrets placement audit: search application configs for backend endpoints and API keys. The blueprint's Action 2 expects them only at the Gateway tier. Direct app-to-internet telemetry egress is the anti-pattern being removed.