Skip to content

How-to Guides

Install, configure, scale, build, and upgrade recipes for the Collector, checked against v1.67.0/v0.161.0 on 2026-09-25. Configuration fragments carry only verified keys. Defaults and stability levels are in the reference. The reasons behind each recipe are in the explanation.

Installation Paths

Three documented routes on Kubernetes, in the order the docs present them:

  1. Pinned manifest, as a starting point. A single command installs the Collector as a DaemonSet (agent) plus one gateway instance (Service + Deployment). The docs call this a starting point.
  2. Helm charts, the named path for production installs.
  3. OpenTelemetry Operator, which can "provision and maintain an OpenTelemetry Collector instance", with automatic upgrade handling, Service objects generated from the OTel configuration, and automatic sidecar injection into deployments.

Whatever the vehicle, remember the config-activation rule from the explanation: components count only when listed under service.

Pinned example manifest (agent DaemonSet plus gateway), from the Kubernetes install page:

kubectl apply -f https://raw.githubusercontent.com/open-telemetry/opentelemetry-collector/v0.161.0/examples/k8s/otel-config.yaml

Helm, using the Kubernetes distribution image as the chart README shows (mode must be daemonset, deployment, or statefulset):

helm repo add open-telemetry https://open-telemetry.github.io/opentelemetry-helm-charts
helm install my-opentelemetry-collector open-telemetry/opentelemetry-collector \
  --set mode=daemonset \
  --set image.repository="ghcr.io/open-telemetry/opentelemetry-collector-releases/opentelemetry-collector-k8s" \
  --set command.name="otelcol-k8s"

Operator (requires cert-manager for the default webhook certificates):

kubectl apply -f https://github.com/open-telemetry/opentelemetry-operator/releases/latest/download/opentelemetry-operator.yaml

Docker, core distribution with a mounted config:

docker run -v "$(pwd)/config.yaml:/etc/otelcol/config.yaml" otel/opentelemetry-collector:0.161.0

Pick the image to match the distribution

otel/opentelemetry-collector is core (otelcol), otel/opentelemetry-collector-contrib is contrib (otelcol-contrib), and the ghcr.io/open-telemetry/opentelemetry-collector-releases/... images mirror all distributions. The binary name inside the image changes with the distribution, which is why the Helm chart needs command.name.

Validate A Config Before Rollout

The Collector binary ships subcommands that catch wiring mistakes without starting pipelines:

otelcol-contrib validate --config=config.yaml            # load + validate, exit non-zero on error
otelcol-contrib components                                # list compiled-in components and their stability
otelcol-contrib print-config --config=config.yaml         # effective config, secrets redacted by default

print-config no longer needs a gate: otelcol.printInitialConfig was stabilized and removed in v0.155.0. Use --mode=unredacted only on a trusted terminal. validate catches ErrSignalNotSupported wiring errors (see signal enforcement), but not components that are configured and never referenced.

Migrate Renamed Component Types

Contrib and core renamed dozens of component types to snake_case between v0.144.0 and v0.160.0 (full table in the reference). Old names still load as deprecated aliases and log a warning, except dynamic_sampling, which became adaptive_tail_sampling with no alias.

  1. Run the new version against your current config and collect the deprecation warnings from the Collector log.
  2. Rename the component keys and every reference under service::pipelines. Keep the /<name> suffix:
exporters:
  otlp_grpc/backend:          # was: otlp/backend
    endpoint: backend:4317
  otlp_http/vendor:           # was: otlphttp/vendor
    endpoint: https://otlp.example.com
receivers:
  file_log:                   # was: filelog
    include: [/var/log/pods/*/*/*.log]
processors:
  k8s_attributes: {}          # was: k8sattributes
service:
  pipelines:
    logs:
      receivers: [file_log]
      processors: [k8s_attributes]
      exporters: [otlp_http/vendor]
  1. Update dashboards and alerts too. Component IDs appear in internal metric attributes, and v0.155.0 renamed memory limiter metrics to otelcol_processor_memory_limiter_*.

Operator and Helm defaults lag

The Operator pins Collector 0.159.0 and the Helm chart ships appVersion 0.160.0. Their generated configs may still use the old names. That keeps working while the aliases exist.

Operator Autoscaling Recipe

Collectors managed by the Operator get built-in HPA. Configure it on the CR rather than creating your own HorizontalPodAutoscaler:

apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata:
  name: gateway
spec:
  mode: deployment            # REQUIRED for autoscaling (or statefulset)
  autoscaler:
    minReplicas: 2
    maxReplicas: 8
    targetCPUUtilization: 90  # the webhook default when no CPU or memory target is set
    targetMemoryUtilization: 80 # illustrative value

Verified constraints:

  • Only works with mode: deployment or statefulset. "HPA only applies to StatefulSets and Deployments in Kubernetes". DaemonSets have no scalable replicas field, so agent collectors cannot be HPA-scaled this way (Operator issue #2605 tracks exactly this). Sidecar-mode collectors are likewise excluded.
  • KEDA's scaling model has the same Deployment/StatefulSet-only restriction.
  • The Operator webhook sets targetCPUUtilization: 90 only when neither a CPU nor a memory target is given.

Scaling Playbook

Distilled doctrine (all verified):

Tier Scale how Load balancing
Agent DaemonSet Vertically, adjust resource limits n/a (host-local)
Gateway without tail sampling Vertical + horizontal Any round-robin LB or K8s Service
Gateway with tail sampling Vertical + horizontal, cautiously load_balancing exporter, routing_key: traceID. Prefer a single well-resourced instance
Prometheus scraping across replicas Horizontal via StatefulSet Target Allocator, not a load balancer (see below)

Cross-reference: the three reference implementations in the guidance topic instantiate these tiers differently at real scale.

Agent-side trace-ID routing into a tail-sampling tier (headless Service otelcol-sampling in namespace observability is an example name):

exporters:
  load_balancing:
    routing_key: traceID
    protocol:
      otlp:
        tls:
          insecure: true      # in-cluster hop; use TLS across trust boundaries
    resolver:
      k8s:
        service: otelcol-sampling.observability
        ports: [4317]

Shard Prometheus Scraping With The Target Allocator

Use this when a prometheus receiver must run on more than one replica without duplicate scrapes. From the Operator's Target Allocator docs:

apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata:
  name: collector-with-ta
spec:
  mode: statefulset
  targetAllocator:
    enabled: true
    prometheusCR:
      enabled: true              # discover ServiceMonitor / PodMonitor CRs
      serviceMonitorSelector: {}
      podMonitorSelector: {}
  config:
    receivers:
      prometheus:
        config:
          scrape_configs:
          - job_name: 'otel-collector'
            scrape_interval: 10s
            static_configs:
            - targets: [ '0.0.0.0:8888' ]
    exporters:
      debug: {}
    service:
      pipelines:
        metrics:
          receivers: [prometheus]
          exporters: [debug]

Checklist:

  • Escape $ as $$ in relabel replacement values inside the Collector config. The Operator turns them back into $ for the TA.
  • With prometheusCR.enabled, install the ServiceMonitor/PodMonitor CRDs and grant the TA ServiceAccount the RBAC the docs list. The Operator does not create those roles for you.
  • For per-node scraping from a DaemonSet, set allocationStrategy: per-node.
  • For more control than spec.targetAllocator offers, create a separate TargetAllocator CR (opentelemetry.io/v1alpha1) and link it with the opentelemetry.io/target-allocator: <name> label.

Custom Builds With ocb

The OpenTelemetry Collector Builder (ocb) generates a complete custom Go binary mixing three component classes:

  1. your own custom components,
  2. upstream core and contrib components,
  3. any other publicly available Go components.

Documented motivations: a smaller binary footprint, or capabilities upstream does not ship (authenticator extensions, receivers, processors, exporters, connectors). The builder version tracks the release train. The README recommends the release binaries or the otel/opentelemetry-collector-builder image over go install.

Minimal manifest pinned to the current train (component modules are v0.161.0, providers are in the stable set at v1.67.0):

dist:
  name: otelcol-custom
  description: Custom OpenTelemetry Collector
  output_path: ./otelcol-custom
receivers:
  - gomod: go.opentelemetry.io/collector/receiver/otlpreceiver v0.161.0
processors:
  - gomod: go.opentelemetry.io/collector/processor/memorylimiterprocessor v0.161.0
  - gomod: go.opentelemetry.io/collector/processor/batchprocessor v0.161.0
  - gomod: github.com/open-telemetry/opentelemetry-collector-contrib/processor/k8sattributesprocessor v0.161.0
exporters:
  - gomod: go.opentelemetry.io/collector/exporter/otlpexporter v0.161.0
  - gomod: go.opentelemetry.io/collector/exporter/debugexporter v0.161.0
extensions:
  - gomod: github.com/open-telemetry/opentelemetry-collector-contrib/extension/healthcheckextension v0.161.0
providers:
  - gomod: go.opentelemetry.io/collector/confmap/provider/envprovider v1.67.0
  - gomod: go.opentelemetry.io/collector/confmap/provider/fileprovider v1.67.0
  - gomod: go.opentelemetry.io/collector/confmap/provider/yamlprovider v1.67.0

Build it with the container image (the output directory must match dist::output_path):

docker run -v "$(pwd)/builder-config.yaml:/build/builder-config.yaml" \
  -v "$(pwd)/output:/build/otelcol-custom" \
  otel/opentelemetry-collector-builder:0.161.0 --config=/build/builder-config.yaml

Or with Go 1.26+ (the minimum since v0.160.0):

go install go.opentelemetry.io/collector/cmd/builder@v0.161.0
builder --config=builder-config.yaml

Split generation and compilation for CI: --skip-compilation to generate and commit sources, then --skip-generate --skip-get-modules to compile only.

Forking Caveat

Forked components that import the collector repos' internal/ packages can fail to compile against different versions. Keep forks shallow or re-implement against public APIs. The builder's strict versioning check rejects manifests whose modules resolve to a different core minor version.

Move Batching Into The Exporter

The batch processor still works (beta), but the project's direction is exporter-side batching (see Batching Migration). Exporter-side batching sits after the queue, so a persistent queue stores the same unit that is sent:

exporters:
  otlp_grpc/backend:
    endpoint: backend:4317
    sending_queue:
      enabled: true
      sizer: requests
      queue_size: 1000
      batch:
        flush_timeout: 200ms
        min_size: 8192          # measured in items, because batch::sizer defaults to items
        max_size: 16384         # 0 = no splitting
service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter]   # batch processor removed
      exporters: [otlp_grpc/backend]

Partition batches by tenant header with sending_queue::batch::partition::metadata_keys (v0.147.0+). Watch otelcol_exporter_queue_batch_send_size and _bytes. Since v0.159.0 they are recorded after batching and only when batch is configured. The queue_batch processor is the drop-in pipeline replacement, but it is development-stage and in no distribution, so build it with ocb if you want it.

Enable Profiles Pipelines

Profiles are alpha. Use them for evaluation, not critical production (the SIG's own guidance).

otelcol-contrib --config=config.yaml --feature-gates=service.profilesSupport
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
exporters:
  otlp_grpc/profiles:
    endpoint: profiles-backend:4317
service:
  pipelines:
    profiles:
      receivers: [otlp]
      exporters: [otlp_grpc/profiles]

Without the gate, a profiles pipeline fails config validation. For whole-host eBPF profiling, run the otelcol-ebpf-profiler distribution rather than adding the profiler to contrib.

Resiliency Recipes

All defaults below are documented and code-cross-checked (see the explanation for mechanics).

Outage-tolerant exporter with persistent queueing and indefinite retry:

extensions:
  file_storage:
    directory: /var/lib/otelcol/storage
    timeout: 1s
    max_size: 10737418240                 # 10 GiB per bbolt file; v0.156.0+; unset = unlimited
    compaction:
      directory: /var/lib/otelcol/storage # keep compaction on the same volume
      on_rebound: true                    # default false; opt-in drain compaction
      # defaults: rebound_needed_threshold_mib 100, trigger 10 MiB, check_interval 5s

exporters:
  otlp_grpc/backend:
    endpoint: backend:4317                # credentials belong in the exporter's auth block
    retry_on_failure:
      enabled: true                       # 5s x1.5 jittered, cap 30s per docs+code
      max_elapsed_time: 0                 # never give up while backend is down
    sending_queue:
      enabled: true
      storage: file_storage               # removes the in-memory queue; at-least-once delivery
      queue_size: 5000                    # 1000 batches is the default; community finds it disk-limiting
      block_on_overflow: false            # false = drop-on-full; true = block until space

service:
  extensions: [file_storage]

Memory-ceiling guard upstream of exporters:

processors:
  memory_limiter:
    check_interval: 1s
    limit_mib: 400        # process RSS typically runs ~50MiB above this
    spike_limit_mib: 80   # soft limit = limit_mib - spike_limit_mib; refusals above it

Operational gotchas, all verified:

  • Watch otelcol_exporter_enqueue_failed_* and otelcol_processor_refused_spans. Refused data relies on upstream components retrying correctly. A misbehaving receiver turns backpressure into loss.
  • Auth-extension context does not survive the persistent queue. Configure exporter-side auth explicitly.
  • Crash windows make persistent queues at-least-once. Assume downstream deduplication if you need exactly-once semantics.
  • When configuring both fixed (limit_mib) and percentage memory limits, fixed silently wins.
  • Tail-sampler sizing: budget RAM as arrival rate × decision_wait (Little's Law: 100k spans/s at a 2-minute wait buffers ~12M spans). Alert on otelcol_processor_tail_sampling_sampling_trace_dropped_too_early, because default overflow silently evicts the oldest traces rather than refusing input. Set block_on_overflow: true (contrib v0.133.0+) when loss is unacceptable.
  • queue_size counts batches, not bytes, under the default sizer. The operative ceiling is real disk. Monitor actual file growth rather than trusting the item budget (an unresolved stall on a 10Gi PVC is documented as contrib #30770).

Move Tail-Sampling Buffers To Disk

Alpha. The pebble_tail_storage extension holds pending traces on disk instead of the heap. It requires a feature gate and cannot be combined with num_shards > 1:

otelcol-contrib --config=config.yaml --feature-gates=processor.tailsamplingprocessor.tailstorageextension
extensions:
  pebble_tail_storage:
    directory: /var/lib/otelcol/pebble-tail-storage
    max_storage_size_mib: 10240          # v0.159.0+
processors:
  tail_sampling:
    decision_wait: 30s
    tail_storage: pebble_tail_storage
    policies:
      - name: errors
        type: status_code
        status_code: {status_codes: [ERROR]}
service:
  extensions: [pebble_tail_storage]

Expect lower heap and higher CPU. The measured trade is summarized in the explanation. The extension drops its database on start (v0.154.0), so it does not preserve pending traces across restarts.

Expose Internal Telemetry

Scrape the Collector's own metrics from other pods (the default binds to localhost:8888):

service:
  telemetry:
    metrics:
      level: normal
      readers:
        - pull:
            exporter:
              prometheus:
                host: 0.0.0.0
                port: 8888
                without_type_suffix: true
                without_units: true

Setting readers yourself drops the defaults that the implicit reader applies, which is why the two without_* options are repeated here. Since v0.149.0 the service_name/service_instance_id/service_version labels are no longer stamped on every internal series. Join on target_info instead. If you scrape this endpoint with a prometheus receiver, the v0.148.0 CHANGELOG "Known Issues" entry gives a metric_relabel_configs mapping back to dotted service.* names (core issue #14814).

Run The OpAMP Supervisor

The Supervisor starts and stops the Collector and relays remote config from an OpAMP server. Minimal Linux config, adapted from the upstream example:

server:
  endpoint: wss://opamp.example.internal:4320/v1/opamp
capabilities:
  reports_effective_config: true
  reports_own_metrics: true
  reports_health: true
  accepts_remote_config: true      # off by default; also enables ReportsRemoteConfig (v0.161.0+)
agent:
  executable: /usr/bin/otelcol-contrib
  automatic_config_rollback: true  # restore last working config if a remote config fails
storage:
  directory: /var/lib/otelcol/supervisor
./opampsupervisor --config=supervisor.yaml

Do not set accepts_packages: true. At v0.161.0 it still aborts startup ("not yet fully implemented"). Upgrade Collector binaries through your normal package or image pipeline.

Upgrade Routine

Every two weeks a new minor lands, and there is no LTS line (release facts in the reference). A routine that keeps upgrades cheap:

  1. Read the core and contrib CHANGELOG sections "Breaking changes" and "Deprecations" for every minor you skip. Renames and removed gates are the usual breakers.
  2. Keep core, contrib, and builder on the same minor. Mixed trains fail the builder's strict versioning check.
  3. Validate the new binary against production configs (validate, then a canary). Watch for deprecation warnings.
  4. Re-check the time-sensitive surfaces: internal-telemetry readers, experimental traces self-telemetry, feature-gated profiles, histogram bucket changes (v0.157.0 changed batch-size buckets), and the distribution list (otelcol-prometheus is the newest entry).
  5. Where the Operator manages Collectors, its automatic upgrade handling moves CRs to the Operator's pinned Collector version unless spec.image is custom.
  6. Critical CVEs (CVSSv3 >= 9.0) get a release within five business days. Other security fixes ship within about 30 days, usually in the regular release.

Check the latest release from a terminal:

curl -s https://api.github.com/repos/open-telemetry/opentelemetry-collector/releases/latest \
  | jq -r '.name'   # e.g. "v1.67.0/v0.161.0"
curl -s https://proxy.golang.org/go.opentelemetry.io/collector/pdata/@latest   # stable-module version and timestamp

Release Hygiene

This section's checklist moved into the Upgrade Routine above. The underlying facts (cadence, dual tagging, security SLA, no LTS) are in the reference.