Skip to content

Explanation

How the Collector engine works and why it is built that way. The internals below were first verified against live docs and source code on 2026-08-27 and re-checked against v1.67.0/v0.161.0 on 2026-09-25. Confidence labels come from the original adversarial verification: "medium" marks a 2-1 split that survived. Look-up tables (defaults, ports, stability levels) are in the reference. Tasks are in the how-to guides.

Architecture Overview

A Collector binary is a service that hosts five component classes. Receivers, processors, and exporters are wired into typed pipelines. Connectors join pipelines. Extensions run beside the pipelines (health, auth, storage, OpAMP) and never touch the data path directly.

The diagram shows a typical gateway build with real component types from core and contrib.

flowchart LR
    subgraph SVC["service (otelcol)"]
        direction LR
        subgraph TP["traces pipeline"]
            OTLPR["otlp receiver<br/>4317 / 4318"] --> ML["memory_limiter"]
            ML --> K8S["k8s_attributes"]
            K8S --> TS["tail_sampling"]
            TS --> FO["fanoutconsumer"]
        end
        FO --> EXG["otlp_grpc exporter"]
        FO --> SM["span_metrics connector"]
        subgraph MP["metrics pipeline"]
            SM --> MB["batch"]
            PROM["prometheus receiver"] --> MB
            MB --> PRW["prometheus_remote_write exporter"]
        end
        subgraph EXT["extensions"]
            HC["health_check"]
            FS["file_storage"]
            OPA["opamp"]
            AUTH["bearertokenauth"]
        end
    end
    EXG -.->|"sending_queue.storage"| FS
    EXG -.->|"auth"| AUTH
    EXG --> BE[("OTLP backend")]
    PRW --> TSDB[("Prometheus-compatible TSDB")]

The same source tree produces every official distribution. The distributions differ only in which component modules the builder compiles in (see Distributions And The Builder).

Pipeline Engine Internals

The simplest pipeline shape, as it exists in code:

flowchart LR
    R1["otlp receiver"] --> P1["memory_limiter"]
    R2["prometheus receiver"] --> P1
    P1 -->|"may drop data<br/>sampling / filtering"| P2["last processor"]
    P2 --> FO["fanoutconsumer"]
    FO -->|"copy if MutatesData"| E1["otlp_grpc exporter"]
    FO -->|"copy if MutatesData"| E2["debug exporter"]

Data flow is push-based and sequential within a pipeline:

  1. All receivers feed the first processor.
  2. Each processor pushes onward and can drop data (this is how sampling and filtering work).
  3. The last processor feeds the built-in fanoutconsumer, so "each exporter gets a copy of each data element." It lives in the internal/fanoutconsumer module (part of the beta module set). The fanout node exists even when only one exporter is configured.
  4. Whether a copy happens depends on component MutatesData capabilities. Mutation-sensitive consumers receive defensive copies.

pdata: The Internal Data Model

pdata ("pipeline data") is the canonical in-memory model for everything the Collector touches:

  • All received data is converted into pdata. It traverses the entire pipeline in that format. Exporters convert out only when sending. Consumer interfaces (ConsumeTraces / ConsumeMetrics / ConsumeLogs, plus ConsumeProfiles in xconsumer) are typed on pdata signals. There is no raw-bytes passthrough component type in signal pipelines.
  • Implementation detail (medium confidence, code-confirmed): pdata wraps OTLP protobuf structs as underlying storage through a private orig pointer, which makes translation to and from the OTLP wire protocol efficient. The representation is deliberately unexported "so that we are free to make changes to it in the future." Since mid-2025 those structs are pdatagen-generated mirrors under pdata/internal, not literal protoc output. The custom proto encoding became unconditional when the pdata.useCustomProtoEncoding gate was removed in v0.153.0, and reference counting (pdata.enableRefCounting) was stabilized in the same release.
  • pdata has been a v1.x stable module since v1.0.0 (2023-11-27). pdata/pprofile, the profiles model, is still in the beta module set and changes almost every release.

Design Consequence

Because pdata's storage is private, components manipulate telemetry through the accessor API rather than raw protobuf access. Upstream can reshape internals without breaking the component ecosystem.

Signal Enforcement And Config Activation

Pipelines are typed by signal: traces, metrics, logs, and profiles. Profiles pipelines are accepted only when the alpha service.profilesSupport feature gate is enabled (see Profiles Signal). If any referenced receiver, processor, or exporter lacks support for its pipeline's type, the Collector fails at configuration load time with pipeline.ErrSignalNotSupported (defined in pipeline/pipeline.go). Fail-fast at startup beats silent drop mid-flight. Treat config-load errors as the contract tests of your pipeline wiring. Nuance: contrib's receivercreator handles this error for dynamic receivers, which is outside static-config scope.

Activation semantics matter operationally:

receivers:
  otlp: {}          # configured...
exporters:
  debug: {}
service:
  pipelines:
    traces:
      receivers: [otlp]      # ...but NOT enabled until referenced here
      exporters: [debug]
  • Configuring any component does nothing until it appears in the service section: extensions directly, receivers/processors/exporters/connectors via pipelines.
  • Connectors must appear on both ends of their joined pipelines. Authenticator extensions must additionally be referenced from auth configuration.
  • Pipelines require at least one receiver and one exporter. Processors are optional (though recommended ones exist). Startup validation enforces this.

Self-observability lives in the nested service.telemetry section with logs and metrics subsections plus an experimental traces option. The old metrics::address key was deprecated in v0.111.0 and has been ignored by default since v0.123.0. Metrics readers replace it:

service:
  telemetry:
    metrics:
      readers:
        - pull:
            exporter:
              prometheus:
                host: localhost
                port: 8888

Startup-validation edge cases worth knowing (all code-confirmed):

  • Connector architectures still satisfy the required exporter slot of a pipeline. The rule counts connectors where they apply.
  • service.AllowNoPipelines (alpha gate) concerns running with zero pipelines. It does not waive the receiver/exporter requirements of pipelines you do declare.
  • Components that load but are never wired remain inert. Silent misconfiguration surfaces as "no data", which is why this got its own callout above.
  • Since v0.156.0 receivers always start after every other component. Before that, receivers sharing an internal implementation (such as the OTLP receiver) could push data into a pipeline whose components had not finished starting.

Config Reload

A config change normally rebuilds the whole service. Two feature gates added in v0.157.0 narrow that: service.partialReload (alpha) together with service.partialReloadReceivers (beta) restart only the receivers when the non-receiver sections are unchanged. This keeps processor state (for example tail-sampling buffers) and exporter queues alive across receiver-only edits. The design is in the core repo's partial-reload RFC.

Deployment Modes

The docs prescribe two modes, which together define the canonical topology:

Mode Shape Role
Agent daemon, sidecar, or DaemonSet. VM binary or container Deployed independently of SDKs. It can aggregate raw measurements for languages lacking in-process stats
Gateway centrally run instances, per cluster or per region Receives from agents and libraries over supported protocols, processes centrally, forwards to configured exporters

The example Kubernetes manifest in the docs deploys both at once: an agent DaemonSet plus one gateway Deployment. Agents collect application and host telemetry and ship to gateways over OTLP/gRPC port 4317 across the internal cluster network. Gateways do centralized processing (filtering, sampling) and TLS egress to backends.

flowchart TB
    subgraph Nodes["Every node"]
        PODS["App pods<br/>OTel SDKs"] --- AGENT["Agent collector<br/>DaemonSet, otelcol-k8s"]
        KUBELET["kubelet"] --- AGENT
    end
    AGENT -->|"OTLP/gRPC 4317<br/>internal cluster network"| LB["load_balancing exporter<br/>or K8s Service"]
    LB --> GW["Gateway Deployment<br/>tail_sampling, filter, TLS egress"]
    GW --> B1[("Self-hosted backend")]
    GW --> B2[("SaaS backend")]

Aspirational Sentence In The Docs

The architecture page contains legacy design language about agents eventually pushing configuration "(such as sampling probability)" down to libraries. Verifiers flagged this as aspirational, not implemented. The real control-plane analog today is OpAMP and the Supervisor managing Collector config, not injecting settings into SDKs. Do not build plans around that sentence.

Tail Sampling Placement

Hard rule from the docs caution block: "The tail-sampling processor can make accurate decisions only if all spans for a trace arrive at the same Collector instance." Hence:

  • Tail sampling runs gateway-side, never distributed across agents.
  • Agents feed it through the load-balancing exporter (type load_balancing since v0.153.0, loadbalancing is a deprecated alias) with routing_key: traceID, the default for traces. That makes affinity sticky per trace rather than per request batch.
  • Sampling decisions downstream stay consistent only when all spans of one trace share an instance (see fanout semantics).
  • Official guidance: prefer a single well-resourced tail-sampling gateway unless you have a sticky-routing strategy that works. Multi-instance setups hit routing re-splitting and decision-cache consistency caveats. The processor README is blunter: all spans of a trace must land on one instance.
  • Scope nuance: "gateways only" applies within multi-instance topologies. A lone single-agent deployment hosts it fine.

Multi-instance routing mechanics. The supported pattern hashes consistently to pick a downstream replica: trace ID by default, service when feeding span-to-metrics pipelines (to avoid label collisions). It is only eventually consistent. The resolvers (static, dns, k8s, aws_cloud_map; DNS refresh default 5s) refresh independently, so during scale-up or scale-down replicas briefly disagree about the backend set. Maintainers advise lowering the resolver interval in highly elastic environments. The k8s resolver converges faster than dns, and since v0.159.0 it skips EndpointSlice endpoints marked not ready. Practitioner reports document harsher-than-documented behavior, including data loss on DNS record change (contrib issue #35378). The exporter does not re-route to a healthy endpoint on failure by default, because its sub-exporters' queue and retry settings start disabled.

Recent processor changes that shift the calculus:

  • num_shards (v0.159.0) runs up to 256 parallel event loops sharded by trace ID. It cannot be combined with tail_storage.
  • sampling_strategy: span-ingest decides per incoming batch instead of after decision_wait, trading policy flexibility (stateful policies are rejected) for lower buffering.
  • The processor.tailsamplingprocessor.usetracestate gate (alpha, v0.157.0) makes the probabilistic policy consume and rewrite W3C tracestate thresholds, so adjusted counts stay correct end to end.
  • Contrib also ships adaptive_tail_sampling (renamed from dynamic_sampling in v0.160.0 with no alias), a separate development-stage processor for adaptive rates.

Resiliency Internals

What happens between exporter and backend, per docs cross-checked against code. Exact defaults are in the reference.

The sequence shows one export attempt through the exporterhelper with a persistent queue.

sequenceDiagram
    participant P as Last processor
    participant Q as sending_queue
    participant S as file_storage (bbolt)
    participant C as queue consumer (x10)
    participant R as retry_on_failure
    participant B as Backend
    P->>Q: ConsumeTraces(batch)
    alt queue full and block_on_overflow false
        Q-->>P: error, otelcol_exporter_enqueue_failed_spans++
    else space available
        Q->>S: write request (WAL)
        Q-->>P: nil (accepted)
        C->>S: read next request
        C->>R: send
        R->>B: export (timeout 5s)
        B-->>R: retryable error
        R->>B: retry after 5s x1.5 jittered, cap 30s
        B-->>R: success
        R->>S: delete request
    end
Mechanism Default behavior Knob
Sending queue capacity Drop-on-full at queue_size=1000 in units of sizer (requests = batches, most performant. bytes least performant) block_on_overflow: true opts into blocking until space or timeout instead of dropping
Queue drain 10 consumers (num_consumers) Per-exporter config
Retry policy Enabled. 5s initial, x1.5 jittered backoff capped at 30s. It gives up per batch after 300s and drops the data (official data-loss circumstance #1) max_elapsed_time: 0 retries indefinitely through outages

Rejected-before-enqueue data never gets to retry logic. Observability lives in otelcol_exporter_enqueue_failed_{spans,metric_points,log_records}, otelcol_exporter_queue_size vs _capacity, throttling logs ("Dropping data because sending_queue is full"), and upstream otelcol_receiver_refused_* movement. Since v0.159.0, otelcol_exporter_enqueue_size and _size_bytes record enqueue-time sizes, and otelcol_exporter_queue_batch_send_size* record post-batching sizes only when sending_queue::batch is configured.

Persistent queues. Pointing sending_queue.storage at a storage extension (file_storage being "a popular and safe choice") removes the in-memory queue entirely. The queue becomes a disk write-ahead log written before each export attempt, and after a kill or crash exports resume from where they stopped. Behavioral consequences worth knowing before enabling:

  • Delivery becomes at-least-once. Duplicates can occur if the process dies between backend success and storage delete.
  • Auth-extension context is not propagated through the persistent queue.
  • The file_storage extension can cap each bbolt database file via max_size (bytes, per component instance, unset = unlimited). Writes that force growth past the cap fail with storage-full errors. The option landed in contrib v0.156.0 (July 2026), so older Collectors lack it.
  • Opt-in online "rebound" compaction exists for outage-then-drain workloads. Enable it via compaction.on_rebound. It triggers only once allocated data first exceeded rebound_needed_threshold_mib (default 100 MiB) and later fell below rebound_trigger_threshold_mib (default 10 MiB), with a check every 5s. With max_size set, both thresholds must fit under it.

memory_limiter refuses incoming telemetry above the soft limit (limit_mib - spike_limit_mib) by returning a non-permanent error to the preceding component, which is expected to retry. This propagates backpressure rather than silently dropping. Surface: otelcol_processor_refused_spans (and signal siblings). Real process RSS typically runs about 50 MiB above limit_mib. Percentage mode (limit_percentage, spike defaulting to 20% of it) is documented for Linux with cgroups. (Medium confidence) When both fixed and percentage settings are present, fixed silently wins. Since v0.156.0 forced GC runs back off exponentially when they reclaim less than 5% while above the soft limit (caps max_gc_interval_when_soft_limited and _hard_limited, default 30s). This fixed a long-standing CPU-burning GC loop when an exporter was stuck (core issue #4981).

Batching Migration

Batching is moving out of the pipeline and into the exporter. The core RFC docs/rfcs/batching-migration.md describes the path:

  • Exporters already batch inside sending_queue (sending_queue::batch). Batching after the queue means one batch per backend request, and the persistent queue sees the same unit it sends.
  • v0.158.0 added the queue_batch processor, built on the same exporterhelper code and configuration. With defaults mapped from the batch processor (block_on_overflow: true, queue_size: 10 requests, batch min_size 8192 items, flush_timeout 200ms), it is the planned replacement for batch. It is still in development and in no distribution.
  • v0.159.0 added the pkg.exporterhelper.queueBatchEnabled gate, which turns batching on in NewDefaultQueueConfig() for exporters built on the helper.
  • The batch processor is still beta and still shipped. Nothing forces a migration yet.

Profiles Signal

Profiles are the fourth OTLP signal. The Profiling SIG announced a public alpha on 2026-03-26, pointing at Collector v0.148.0 or newer. In the Collector:

  • service.profilesSupport (alpha since v0.112.0) must be enabled before a profiles pipeline is accepted.
  • Since v0.149.0 the core OTLP receiver, otlp_grpc and otlp_http exporters, debug, nop, forward, and pdata/pprofile are alpha for profiles. Contrib adds the pprof receiver (alpha) and development-stage profiles support in k8s_attributes, transform, filter, resource_detection, and kafka.
  • The eBPF profiler donated by Elastic runs as a Collector receiver and ships as the otelcol-ebpf-profiler distribution (release artifacts since v0.133.0).
  • The SIG says the signal "should not be used for critical production workloads". pprofile APIs are still being removed and reshaped (for example v0.161.0 removed deprecated AggregationTemporality and Duration fields).

OpAMP Management Plane

flowchart LR
    SRV["OpAMP server"] <-->|"remote config, status, heartbeat"| SUP["opampsupervisor"]
    SUP -->|"start, stop, SIGHUP"| COL["Collector process"]
    COL <-->|"local OpAMP"| EXTN["opamp extension"]
    EXTN --- SUP
    CFG[("effective.yaml<br/>merged local + remote")] -.-> COL
    LAST[("last received and<br/>last working config")] -.-> SUP

Split verdict, re-verified against v0.161.0:

  • Spec status: OpAMP is formally Beta (opamp-spec v0.20.0). Breaking changes remain possible between releases, with no spec-level production-readiness guarantee. Individual message fields carry their own status, and several newer ones (connection settings requests, custom capabilities, available components) are still Development. Core functions: remote configuration delivery, agent status reporting, heartbeats (recommended default interval 30s).
  • Remote configuration: production-viable but conservative in the contrib Supervisor (component stability alpha). accepts_remote_config is off by default and, since v0.161.0, also enables ReportsRemoteConfig. Applying remote config merges it with optional local config and restarts the Collector process. An opt-in agent::use_hup_config_reload reloads via SIGHUP instead (non-Windows only, introduced around v0.130.0). Receiving an empty config map stops the Collector until a non-empty one arrives. Outage fallback keeps the last persisted config running across Supervisor restarts while it reconnects with exponential backoff. Newer safety nets: a startup fallback config until the first successful connection (v0.154.0) and agent::automatic_config_rollback to the last working config when a remote config fails to start (v0.157.0, default off).
  • Executable and package upgrades: not production-ready upstream. accepts_packages parses in Supervisor config, but enabling it still blocks startup with "accepts_packages capability is not yet fully implemented" at v0.161.0. Groundwork landed in v0.159.0 (agent_binary configuration, tar.gz archives), and v0.158.0 folded reports_package_statuses into accepts_packages. Tracked in open issues #47272 (packages) and #33947 (binary update).

Refuted Overclaim

Verifiers rejected (0-3) any claim that the OpAMP spec guarantees secure auto-update including downgrades. Weigh vendor-blog claims of guaranteed safe fleet self-updates against the Beta spec status and the still-blocked package-management implementation.

Target Allocator

Prometheus scraping does not shard itself. If three Collector replicas each run the same prometheus receiver config, each scrapes every target and the backend gets three copies. The Operator's Target Allocator (TA) solves this by moving discovery out of the Collectors.

sequenceDiagram
    participant TA as Target Allocator
    participant K as Kubernetes API
    participant C as Collector replicas (prometheus receiver)
    participant T as Scrape targets
    TA->>K: watch ServiceMonitor, PodMonitor, scrape configs
    TA->>K: watch Collector pods
    TA->>TA: assign targets (consistent-hashing by default)
    C->>TA: HTTP SD request with collector_id
    TA-->>C: this replica's targets
    C->>T: scrape

Why it is shaped this way:

  • The Operator rewrites the Collector's prometheus receiver to point at the TA (target_allocator.endpoint, interval: 30s, collector_id: $POD_NAME). The scrape jobs themselves move to the TA's config.
  • consistent-hashing hashes only the target URL, so label changes do not move targets. Targets still rebalance when the replica count changes. least-weighted trades evenness for stability. per-node pins targets to the Collector on the same node and only makes sense with a DaemonSet.
  • Prometheus Operator CRs are discovered without running Prometheus, but the ServiceMonitor/PodMonitor CRDs must be installed.
  • Scrape credentials from monitors are served to Collectors over mTLS (cert-manager by default, user-provided Secrets since Operator v0.157.0). An allowInsecureAuthSecrets opt-out exists for plain HTTP.

CR fields and Operator versions are in the sibling OpenTelemetry reference. TA facts are in this topic's reference.

Auth And Security Posture

The stable config/configauth and extension/extensionauth modules define two directional authenticator categories, exposed through auth: blocks on confighttp/configgrpc next to their tls: blocks. Server authenticators guard receivers. Client authenticators decorate exporter requests. The full extension matrix is in the reference.

  • basicauth and bearertokenauth are dual-role. The OIDC extension (type oidc) is server-only. oauth2client, sigv4auth, and asapclient are client-only.
  • headers_setter is not an authenticator. It mutates headers rather than implementing either interface.
  • Caveat: the documented extension lists are manually maintained. Check them against the contrib state for each release. An earlier claim that sigv4auth carries a deprecation notice is not supported by its README or metadata.yaml at v0.161.0 (beta, no notice).
  • Receivers and extensions default to localhost endpoints (the Prometheus internal-telemetry endpoint since v0.111.0). Exposing them is an explicit decision, which is the core of the upstream configuration security best practices.
  • Contrib CVEs in 2026 were component-local (GitHub receiver header enforcement, Sentry exporter URL path traversal). A custom build with fewer components has a smaller attack surface. See the reference.

This pairs directly with Skyscanner's and Mastodon's gateway placements in the guidance topic.

Distributions And The Builder

Every official distribution is an ocb build of a manifest.yaml in opentelemetry-collector-releases. The design choice is that the project ships a small set of curated builds and expects production users to build their own:

  • otelcol (core) is a minimal foundation set. otelcol-contrib compiles in nearly everything, including vendor exporters. otelcol-k8s and otelcol-otlp are curated subsets. otelcol-ebpf-profiler bundles the eBPF profiler receiver.
  • otelcol-prometheus (listed in the releases repo in 2026) embeds Prometheus exporters such as blackbox, postgres, stackdriver, and yace as in-process receivers. Its README calls it experimental and "not recommended for production".
  • The builder enforces strict versioning: the core library version the Go toolchain resolves must match the manifest's major and minor versions. That is why mixing components from different release trains fails early.
  • Builds default to CGO_ENABLED=0 and stripped symbols, matching official releases.

Stability Model

Two independent axes describe "how stable" a piece of the Collector is:

  1. Go module version. Modules in the stable set are v1.x and follow semver for their Go API. Everything else is v0.x, including the otelcol and service packages and every component. This axis protects people who write components or embed the Collector.
  2. Component stability level per signal. Development, alpha, beta, stable, plus deprecated and unmaintained. This axis protects people who write YAML. A component can be v1.x only when it is stable for at least one signal. Once it is, its Go API cannot break for any signal, even signals that are still alpha.

This explains why "Collector 1.0" is not a single event. Many APIs are already v1 and the OTLP receiver and exporters are stable for traces, metrics, and logs, yet the binary and most components remain v0. The public roadmap to v1 targets stable configuration and APIs first. As of 2026-09-25 no v1 date is published.

The snake_case rename campaign (filelog to file_log, otlp exporter to otlp_grpc, and dozens more across v0.144.0 to v0.160.0) is part of the same configuration-stabilization effort. Renames ship with deprecated aliases (added via mdatagen's deprecated_type in v0.148.0), so old configs keep loading with a warning. The full list is in the reference.

Mapping Engine Features To The Reference Implementations

The organizational topologies reduce to these engine-level primitives:

Reference implementation Engine primitive they use
Adobe: immutable sidecar config Config-load-time validation plus sidecar mode. Config changes routed to a Deployment collector instead
Adobe: backend choice via header Routing through connectors keyed on OTLP transport metadata
Mastodon: one CR per namespace Operator OpenTelemetryCollector CR in deployment mode, single-pipeline simplicity
Skyscanner: bulk-processing gateways Gateway mode plus fanout exporter copies. Agents limited to scrape duties
All three: OTLP everywhere pdata/OTLP affinity. OTLP is zero-translation into pdata's wrapped storage

That last row generalizes into a design heuristic: components that translate formats belong at pipeline edges (receivers and exporters). This keeps the interior in pdata, so fanout copies stay cheap and signal support stays declaratively verifiable at config load.

Scaling Doctrine

Tiered, as documented:

  • Agents "typically do not require horizontal scaling because they run on each host". Scale them vertically via resource limits. (Verifiers kept the word "typically".)
  • Gateways scale both vertically and horizontally.
  • Load-balancer selection follows sampling needs: a plain round-robin LB or K8s Service without tail sampling, trace-ID-aware routing with it.
  • Automatic horizontal scaling means a Kubernetes HPA on CPU or memory, delivered natively by the Operator for CR-managed collectors in deployment or statefulset mode. The table form and the CRD recipe are in the how-to guides.

Benchmarks And Measured Numbers

Outcome of the numbers-focused pass (2026-08-27): the official measurement corpus is verifiably absent for all three open questions, and every surviving number comes from external labs or issues with stated caveats. Not re-run on 2026-09-25.

Verified Official Absence

Parsing the complete dataset behind the Collector load-test dashboard (data.js, regenerated twice daily, snapshot verified 2026-08-27T04:48Z) yielded exactly 100 CI runs over 59 scenarios / 177 chart labels with zero matches for tail sampling, batch-processor tuning, or persistent-queue sizing. The only "batch" series measure receiver-side TCP write batching, not the batch processor. A code audit agreed: the testbed trace suite wires only batch/attributes/memory_limiter, and testbed/ contained zero tailsampling references. The sending-queue scenarios assert behavior only (a "queue full" log line appears, queued items are retried) and publish no numeric rows. Falsifiable form: one future BenchmarkTailSampling* entry refutes this claim.

Tail-Sampling Memory Scaling Law

Practitioner-derived (issue #31498, Feb 2024, unversioned, consistent with Little's Law, corroborated directionally by Elastic's buffering findings):

buffered spans ≈ arrival rate × decision_wait
for example: 100k spans/s × 2 min decision_wait ≈ 12 million spans held in RAM

Decision latency defaults to roughly one full decision_wait (30s) per trace. The decision_wait_after_root_received opt-in short-circuits this (default 0s). At capacity (num_traces, default 50000) the sampler does not refuse ingest or disable policies. It evicts the oldest buffered trace from its circular buffer before sampling (metric otelcol_processor_tail_sampling_sampling_trace_dropped_too_early, which the README FAQ calls "likely a load issue"). True backpressure requires opt-in block_on_overflow: true (option added in contrib v0.133.0). Both LRU decision caches (decision_cache::sampled_cache_size, non_sampled_cache_size) default to 0, meaning disabled, with only relational sizing advice in the docs and no byte accounting anywhere.

Disk-Backed Decision Storage: The Measured Trade

Pebble-backed trace storage moves the buffer off-heap. It now ships as the pebble_tail_storage extension (first code v0.151.0, alpha since v0.154.0, max_storage_size_mib cap since v0.159.0), plugged into the processor through tail_storage behind the processor.tailsamplingprocessor.tailstorageextension gate. The two credible number sets disagree because they measure different bundles:

Study Scope under test Memory result CPU result
Elastic Observability Labs (2026-07-21) span-ingest strategy plus pebble storage (confounded bundle) peak heap −65.4%, working set −62.9%, RSS −51.7% aggregate ~2×
contrib #42326 PoC (relayed by VictoriaMetrics' KubeCon EU 2026 write-up) pebble-only (isolates storage effect) peak −81.3% / avg −82.2% avg ~7.5×, peak ~5.5×

Neither pins the exact binary version under test or full hardware specs. The magnitudes are compatible given the bundling difference, but treat both as directional. A further low-confidence single-run observation (Meth, IBM, bot-gated link): tail sampling can even increase network egress in that demo, because each exported trace line duplicated source metadata.

Persistent-Queue Sizing Reality

queue_size bounds batches, not bytes, under the default requests sizer, regardless of backing store. Byte accounting exists only through the slowest sizer (bytes). As a result, the operative ceiling is real disk, and it can misbehave independently of the item budget. Contrib issue #30770 (closed without a documented fix; checked 2026-09-28) reports a persistent queue that stopped at about 1,500 batches on a 10Gi PVC with no space left on device, even with queue_size at 1,000,000. The reporter says doubling the PVC did not help. The reported config put compaction under /tmp/, outside the volume. Verifiers declined to generalize either way from it. Two settings bound where and how much the queue writes: the file_storage max_size cap (v0.156.0) and an explicit compaction.directory on the same volume. Nobody in the #30770 thread reports trying either.

Open Gaps

Open questions for this topic are tracked in one place: the topic index.