LGTM Stack Explanation¶
How the Grafana LGTM stack works: the shared design of Loki, Tempo, Mimir and Pyroscope, the 2025-2026 shift to Kafka-decoupled write paths, how signals are stored and correlated, and why the stack is cheap to run but demanding to operate. Look-up tables (versions, ports, config keys, limits, scale figures) live in Reference; tasks and recipes live in How-to Guides. Grafana's own internals (plugins, alerting, dashboards) are covered in Grafana Explanation.
Design Principles¶
Four ideas recur across every LGTM backend:
- One binary, many roles. Mimir, Loki, Tempo and Pyroscope each compile all components into a single binary. The
-targetflag selects which components a process runs (all,distributor,querier, ...). The same code therefore runs as a laptop-sized monolith or as dozens of independently scaled microservices. - Object storage is the database. Long-term data lives in S3, GCS, Azure Blob Storage (or Swift / S3-compatible stores such as MinIO). Local disks hold only write-ahead logs, caches and the most recent data. Object storage is cheap, durable and effectively unbounded, so retention becomes a cost setting rather than a capacity plan.
- Index as little as possible. Mimir indexes labels (as Prometheus does). Loki indexes only stream labels, never log content. Tempo keeps no per-span index and instead relies on trace-ID sharding, bloom filters and the columnar Apache Parquet format. Pyroscope indexes by labels and profile type. Less indexing means cheaper ingestion and storage, paid for with more work at query time.
- Multi-tenant by construction. Every request carries a tenant ID in the
X-Scope-OrgIDheader. Data, limits and caches are partitioned by tenant. Authentication is deliberately left to a gateway in front.
Architecture Overview¶
The diagram below shows the production data flow: Alloy (or an OpenTelemetry Collector) receives and routes all signals, each backend writes to its own bucket, and Grafana queries each backend in its own language.
flowchart LR
subgraph Sources["Instrumented workloads"]
SDK["OTel SDKs /<br/>auto-instrumentation"]
Scrape["Prometheus<br/>exporters"]
Files["Container<br/>log files"]
EBPF["eBPF profilers /<br/>Beyla-OBI"]
end
Alloy["Grafana Alloy<br/>(DaemonSet + clustered scrapers)"]
subgraph Backends["LGTM backends"]
Mimir["Mimir 3.x<br/>metrics"]
Loki["Loki 3.x<br/>logs"]
Tempo["Tempo 3.x<br/>traces"]
Pyro["Pyroscope 2.x<br/>profiles"]
end
Kafka[("Kafka-compatible log<br/>(Mimir ingest storage,<br/>Tempo microservices)")]
subgraph Buckets["Object storage (one bucket each)"]
BM["mimir-blocks"]
BL["loki-chunks"]
BT["tempo-traces"]
BP["pyroscope-data"]
end
Grafana["Grafana<br/>Explore, dashboards,<br/>Drilldown apps, alerting"]
SDK -->|OTLP| Alloy
Scrape -->|scrape| Alloy
Files -->|tail| Alloy
EBPF -->|pprof / OTLP| Alloy
Alloy -->|"remote_write or OTLP"| Mimir
Alloy -->|"OTLP /otlp/v1/logs"| Loki
Alloy -->|OTLP gRPC| Tempo
Alloy -->|push| Pyro
Mimir <-->|produce / consume| Kafka
Tempo <-->|produce / consume| Kafka
Mimir --> BM
Loki --> BL
Tempo --> BT
Pyro --> BP
Tempo -.->|"metrics-generator<br/>remote_write"| Mimir
Grafana -.->|PromQL| Mimir
Grafana -.->|LogQL| Loki
Grafana -.->|TraceQL| Tempo
Grafana -.->|profile queries| Pyro
The Four Pillars and Collection¶
| Pillar | Component | What it stores | Key insight |
|---|---|---|---|
| Metrics | Mimir | Per-tenant Prometheus TSDB blocks | Horizontally scalable, long-term Prometheus; Cortex lineage |
| Logs | Loki | Label-indexed log streams in compressed chunks | Indexes labels only, not content, so it costs a fraction of full-text engines |
| Traces | Tempo | Spans in Apache Parquet blocks | No span index; trace-ID lookup + columnar scans (TraceQL) |
| Profiles | Pyroscope | pprof-derived profile segments and blocks | Links CPU/memory hotspots to code and, via span IDs, to traces |
| Collection | Alloy | Nothing (pipeline agent) | Grafana's OpenTelemetry Collector distribution with Prometheus, Loki and Pyroscope components |
Deployment Modes¶
All backends share the same three-tier philosophy, but 2025-2026 releases removed the middle tier from most of them.
| Mode | What runs | Status in 2026 |
|---|---|---|
| Monolithic | Every component in one process (-target=all) |
Supported everywhere; the starting point for dev, PoC and small production |
| Simple Scalable (SSD) | read, write, backend targets |
Loki: deprecated, removal planned for Loki 4.0; Tempo: scalable-single-binary removed in 3.0; Mimir: experimental read-write mode removed in 3.0 |
| HA monolithic | Several -target=all replicas sharing object storage and a memberlist ring |
Loki's documented replacement for SSD (RF=3, at least three replicas, one main compactor); Mimir allows scaled monolithic; Tempo does not (backend scheduler is a singleton) |
| Microservices | One component per process, scaled independently | Recommended for production at scale; Tempo 3 and Mimir ingest storage add Kafka here |
The reasons for dropping SSD are operational: a read or backend pool still mixes components with very different resource profiles, and the new Kafka-based designs already give the read/write separation that SSD was meant to provide.
The flowchart below summarises how to pick a mode for a new deployment.
flowchart TD
Start["New LGTM deployment"] --> Prod{"Production<br/>traffic?"}
Prod -->|No| Mono["Monolithic per backend<br/>(or grafana/otel-lgtm image)"]
Prod -->|Yes| HA{"Need HA but<br/>small volume?"}
HA -->|Yes| HAM["Loki HA monolithic (RF=3)<br/>Mimir scaled monolithic<br/>Tempo monolithic (single instance)"]
HA -->|No| Kafka{"Can you run a<br/>Kafka-compatible system?"}
Kafka -->|Yes| MS["Microservices:<br/>Mimir ingest storage<br/>Tempo 3 with Kafka<br/>Loki distributed"]
Kafka -->|No| Classic["Mimir classic architecture<br/>Loki distributed<br/>Tempo stays monolithic"]
HAM --> Grow{"Outgrowing<br/>guidelines?"}
Grow -->|Yes| Kafka
Limits of Monolithic Mode at Medium Scale¶
At roughly 1M active series and 100 GB/day of logs, monolithic deployments start to hurt:
- No independent scaling. A heavy query competes with ingestion for CPU and memory in the same process.
- Guidelines are exceeded. Loki documents about 20 GB/day for monolithic; Tempo documents 25-35 MB/s (55k-80k spans/s) for its single supported monolithic instance.
- Memory. A Mimir monolith at 1M series needs tens of GB of RAM (community rule of thumb 32-64 GB) plus fast disk for the TSDB head and store-gateway index headers.
- Availability. A single replica is a single point of failure. Loki HA monolithic and scaled Mimir monoliths help; Tempo 3 needs microservices for HA.
Plan the move to microservices before reaching these limits; because the data lives in object storage, you can redeploy under a different mode with little or no data migration (Tempo 2.x to 3.0 is the exception, see below).
Kafka-Decoupled Write Paths¶
The largest architectural change of 2025-2026 is that Mimir and Tempo moved their durable write path from replicated, stateful ingesters to a Kafka-compatible log (Apache Kafka, Redpanda, WarpStream, and similar). Loki is experimenting with the same idea on its main branch.
Why Kafka¶
In the classic design, distributors replicate each write to three ingesters (RF=3) and wait for a quorum. Ingesters then hold data in memory and a local WAL, serve recent queries, and flush blocks to object storage. Consequences:
- Heavy queries on ingesters can slow down or break ingestion.
- Every sample or span is stored three times until compaction deduplicates it.
- Ingester rollouts and scale-downs are delicate because each pod owns unflushed data.
With a Kafka log in front, the write path ends when Kafka acknowledges the record. Everything downstream consumes asynchronously and can be restarted, replayed or scaled independently.
The sequence below shows the Tempo 3 microservices write and read path; Mimir ingest storage follows the same shape with ingesters in place of live-stores and block-builders.
sequenceDiagram
participant C as Alloy / OTel Collector
participant D as Distributor
participant K as Kafka topic
participant LS as Live-store
participant BB as Block-builder
participant MG as Metrics-generator
participant S3 as Object storage
participant QF as Query-frontend
participant Q as Querier
C->>D: OTLP export (spans)
D->>D: validate, apply tenant limits, shard by trace ID
D->>K: produce records to partitions
K-->>D: ack (durable)
D-->>C: 200 OK
K-->>LS: consume (recent data, WAL)
K-->>BB: consume and build Parquet blocks
BB->>S3: flush block (built once, RF1)
K-->>MG: consume, derive RED metrics
QF->>Q: shard query into jobs
Q->>LS: recent window (about 30-60 min)
Q->>S3: older blocks
Q-->>QF: partial results
Mimir Ingest Storage (3.0+)¶
Since Mimir 3.0 (GA November 2025) the ingest storage architecture is stable and the preferred architecture; the classic architecture remains supported.
- Distributors shard each series to one Kafka partition and acknowledge the client only after Kafka persists the batch.
- Each ingester consumes exactly one partition; ingesters in different zones consume the same partition for read-path HA. The partition number comes from the instance ID suffix (
-<N>). - Each ingester writes its own blocks; the compactor merges and deduplicates them.
- The
mimir-distributedHelm chart (6.0+) turns ingest storage on by default and ships a single-node Kafka for demos; production needs an external Kafka-compatible cluster. - Mimir 3.0 also made the streaming Mimir Query Engine (MQE) the default (Grafana reports up to 92% lower peak memory than the Prometheus engine), made the query-scheduler mandatory, and removed Redis caching.
Tempo 3.0 (May 2026)¶
Tempo 3.0 replaced ingesters and compactors entirely:
| Tempo 2.x | Tempo 3.x (microservices) |
|---|---|
| Distributor replicates to 3 ingesters over gRPC | Distributor writes to Kafka; data stored once (RF1) |
| Ingesters hold traces, cut blocks, flush | Block-builders consume Kafka and write Parquet blocks |
| Ingesters serve recent queries | Live-stores consume Kafka and serve the recent window; zone-aware for HA |
| Compactor | Backend-scheduler + backend-worker (compaction, retention, redaction jobs) |
scalable-single-binary target |
Removed |
Monolithic Tempo 3 does not need Kafka: the distributor pushes in-process to the live-store and metrics-generator, and the live-store flushes blocks itself. Microservices mode requires Kafka. There is no in-place downgrade from 3.0 to 2.x; microservices users migrate by running 3.0 in parallel and switching traffic.
Trade-off: the design lowers storage and replication cost and isolates reads from writes, but adds a Kafka cluster to operate (or pay for), and recent-data queries now fail fast when a live-store lags (live_store.fail_on_high_lag: true) rather than returning incomplete results.
Loki and Kafka¶
Loki 3.x still uses the classic distributor-ingester path in all supported modes. Its main branch is building a Kafka-fed data object (dataobj) storage format with a new query engine and dataobj compactor; these are experimental and not a supported deployment mode as of Loki 3.7.
Pyroscope v2 (April 2026)¶
Pyroscope 2.0 made the v2 storage architecture the default. It takes a different route from Mimir and Tempo: no Kafka, and no ingesters.
- Segment-writers (stateless, diskless) accumulate profiles for a short window, co-locate profiles from the same service, and write one object per shard straight to object storage. By default, ingestion is synchronous: the client is acknowledged only once the segment is stored and indexed (median under 500 ms).
- The metastore is the only stateful component. It keeps the metadata index of all objects and coordinates compaction, replicated with Raft.
- Compaction-workers merge segments into blocks; query-backends read object storage directly and can scale out quickly.
- v1 (ingester-based) remains selectable with
-write-path=ingester/-architecture.storage=v1; the default-architecture.storagevaluev1-v2-dualreads both during migration.
Storage Data Models¶
Mimir: Per-Tenant TSDB Blocks¶
Mimir stores each tenant's series in its own Prometheus TSDB. Ingesters cut two-hour blocks, each with an index (label postings), metadata (meta.json) and chunk files (about 120 samples per chunk). The compactor merges blocks into larger ranges (12 h, 24 h by default), deduplicates replicas, and can split-and-merge for very large tenants. Store-gateways lazily load block index headers to serve historical queries.
Loki: Streams, Chunks and Structured Metadata¶
A Loki stream is the unique combination of a tenant and its label set. Each stream's lines are compressed into chunks and indexed by the TSDB index (schema v13). Three places hold per-line data:
| Where | Indexed? | Use for |
|---|---|---|
| Stream labels | Yes (defines the stream) | Low-cardinality context: namespace, cluster, service_name, env |
| Structured metadata | No (stored alongside each line) | High-cardinality context: trace_id, pod, user_id, k8s.pod.uid |
| Log line | No | The message itself; filtered with line filters and parsers at query time |
Structured metadata (Loki 3.0+, on by default with schema v13) exists because putting high-cardinality values in labels explodes the number of streams and chunks. Loki's native OTLP endpoint maps a configurable set of OTel resource attributes to labels and stores the rest as structured metadata. Experimental bloom filters (Loki 3.3+) can accelerate "needle in a haystack" queries on structured metadata for very large tenants.
Tempo: Parquet Blocks¶
Tempo writes traces into Apache Parquet blocks (vParquet4 in 3.0; vParquet5 is the default for new blocks in 3.1). Each span attribute lives in a column, so a TraceQL query reads only the columns it needs. Dedicated columns promote frequently queried attributes out of generic key-value maps for speed; vParquet5 adds dedicated integer and event columns and low-resolution timestamp columns for faster TraceQL metrics. Each block carries bloom filters for trace-ID lookups.
Pyroscope: Segments and Blocks¶
Profiles are split into symbols, stack traces and samples. Segment-writers store them per shard in segments; compaction-workers merge segments into larger blocks per tenant and service, and the metastore indexes which object holds which time range and labels.
Shared Infrastructure¶
Hash Rings¶
Mimir, Loki, Tempo and Pyroscope use consistent hash rings to assign work (series, streams, traces, profiles, compaction jobs) to instances. The ring state is gossiped with memberlist by default; Consul and etcd remain options. Mimir ingest storage and Tempo 3 add a partition ring that maps Kafka partitions to consumers.
Caching¶
Queries against object storage are only fast with caches in front. Memcached is the standard backend for results, chunks, index and metadata caches; Mimir 3.0 removed Redis, and Tempo 3.1 reworked its experimental Redis cache around Redis Cluster. Without caches, every query reads object storage directly, which is slow and incurs request charges.
Cross-Signal Correlation¶
Moving between signals without copy-pasting IDs is the main reason to run the whole stack. It needs both instrumentation (trace context in every signal) and Grafana data source configuration (links between data sources).
flowchart LR
Metrics["Metrics<br/>(Mimir)"]
Logs["Logs<br/>(Loki)"]
Traces["Traces<br/>(Tempo)"]
Profiles["Profiles<br/>(Pyroscope)"]
Metrics -->|"Exemplars<br/>(trace ID on data point)"| Traces
Traces -->|"Trace to logs<br/>(span attrs to LogQL)"| Logs
Logs -->|"Derived fields /<br/>structured metadata trace_id"| Traces
Traces -->|"Trace to metrics<br/>(span attrs to PromQL)"| Metrics
Traces -->|"Trace to profiles<br/>(span ID labels)"| Profiles
Traces -.->|"metrics-generator<br/>span metrics + service graph"| Metrics
| Link | From to | Mechanism | Where configured |
|---|---|---|---|
| Exemplars | Metrics to traces | Trace IDs attached to samples | App instrumentation + Mimir max_global_exemplars_per_user > 0 + Prometheus data source exemplarTraceIdDestinations |
| Trace to logs | Traces to logs | Span attributes become a Loki query | Tempo data source tracesToLogsV2 |
| Derived fields | Logs to traces | Regex or structured-metadata field yields a trace ID | Loki data source derivedFields |
| Trace to metrics | Traces to metrics | Span attributes become PromQL filters | Tempo data source tracesToMetrics |
| Trace to profiles | Traces to profiles | Span ID labels on profiles (span profiling) | Tempo data source tracesToProfiles + SDK span-profiles integration |
| Span metrics / service graph | Traces to metrics (derived) | Tempo metrics-generator computes RED metrics and edges | Tempo metrics_generator + remote_write to Mimir |
The provisioning file that wires these links is in How-to Guides.
Example Incident Workflow¶
The sequence shows a typical on-call path that uses every correlation link once.
sequenceDiagram
participant SRE as On-call engineer
participant Metrics as Mimir (metrics)
participant Traces as Tempo (traces)
participant Logs as Loki (logs)
participant Profiles as Pyroscope (profiles)
SRE->>Metrics: Alert fires, error_rate above 5%
SRE->>Metrics: Open dashboard, see spike
SRE->>Metrics: Click exemplar on spike
Metrics-->>Traces: Jump to trace ID abc123
SRE->>Traces: See slow span in payment-service (2.3s)
SRE->>Traces: Click Trace to logs
Traces-->>Logs: Query service_name=payment, trace_id=abc123
SRE->>Logs: See connection timeout to DB
SRE->>Traces: Click Trace to profiles
Traces-->>Profiles: Flame graph for the span
SRE->>Profiles: CPU hotspot in connection-pool retry loop
SRE->>SRE: Root cause, DB connection pool exhausted
Query Languages¶
The stack uses four query models that share PromQL's label-matching style. Exact syntax references are in Reference.
PromQL (Metrics, Mimir)¶
# Rate of HTTP requests over 5 minutes, grouped by status code
sum(rate(http_requests_total{job="api"}[5m])) by (status)
# 99th percentile latency
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))
# Alert expression: error rate above 5%
sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) > 0.05
LogQL (Logs, Loki)¶
# Filter error logs from production
{app="payment-service", env="prod"} |= "error" != "timeout"
# Parse JSON logs and filter by status
{app="api"} | json | status >= 500
# Filter on structured metadata without parsing
{service_name="checkout"} | trace_id="abc123"
# Count error logs per minute (metric query)
sum(rate({app="api"} |= "error" [1m])) by (pod)
TraceQL (Traces, Tempo)¶
# Traces with HTTP 500 errors
{span.http.status_code = 500}
# Slow spans in a specific service
{resource.service.name = "checkout" && duration > 2s}
# Frontend spans that have a payment descendant
{resource.service.name = "frontend"} >> {resource.service.name = "payment"}
# TraceQL metrics (GA in Tempo 3.0): p95 latency by service
{span.http.route = "/pay"} | quantile_over_time(duration, .95) by (resource.service.name)
Profile Queries (Pyroscope)¶
Pyroscope selects a profile type and filters it with Prometheus-style label matchers. Most users never type these queries; the Profiles Drilldown app in Grafana is queryless.
# CPU profile for one service
process_cpu:cpu:nanoseconds:cpu:nanoseconds{service_name="payment-service"}
# Allocated memory for one environment
memory:alloc_space:bytes:space:bytes{service_name="api", env="production"}
Multi-Tenancy and Security Model¶
Tenant Isolation¶
All backends use the X-Scope-OrgID header as the tenant identifier. Mimir and Loki enable multi-tenancy by default; Tempo and Pyroscope default to a single implicit tenant. Each backend stores tenant data under separate prefixes (Mimir: <bucket>/<tenant>/; Loki chunks and index per tenant; Tempo blocks per tenant) and applies per-tenant limits, so tenants cannot read each other's data unless cross-tenant federation is explicitly enabled.
The diagram shows the trust boundary: the backends trust whatever tenant header they receive, so identity must be established before the header is set.
flowchart LR
subgraph Clients["Untrusted zone"]
A1["Alloy (team alpha)"]
A2["Alloy (team beta)"]
U["Grafana users"]
end
GW["Auth gateway<br/>(NGINX / Envoy / GEM-GEL-GET gateway)<br/>validates identity, sets X-Scope-OrgID"]
subgraph Private["Private network"]
M["Mimir"]
L["Loki"]
T["Tempo"]
P["Pyroscope"]
end
S3[("Object storage<br/>SSE-KMS, private endpoint")]
A1 -->|"token / mTLS"| GW
A2 -->|"token / mTLS"| GW
U -->|"SSO session"| GW
GW -->|"X-Scope-OrgID: alpha or beta"| M
GW --> L
GW --> T
GW --> P
M --> S3
L --> S3
T --> S3
P --> S3
Threat Model¶
| Threat | Why it matters in LGTM | Mitigation |
|---|---|---|
| Tenant spoofing | Backends accept any X-Scope-OrgID; open-source editions have no built-in authentication |
Gateway strips client headers and sets the tenant from verified identity |
| Noisy neighbour / DoS | One tenant's cardinality explosion can exhaust ingesters or queriers | Per-tenant ingestion, series/stream/trace and query limits |
| Data exfiltration from storage | All history sits in buckets | Separate bucket per component, least-privilege IAM, SSE-KMS, private endpoints |
| Cross-tenant data mutation | Admin APIs (deletes, redaction) act on tenant data | Tempo 3.1 takes the redaction tenant only from X-Scope-OrgID, ignoring request bodies; restrict admin endpoints at the gateway |
| Credential leakage in logs | Query logs can capture headers | Mimir 3.2 refuses to log credential-bearing headers (Authorization, Cookie, X-Api-Key) |
| Sensitive data in telemetry | Traces and logs often carry PII | Filter or redact in Alloy; Tempo supports on-demand redaction jobs (tempo-cli redact) |
| Supply chain | Images pulled into clusters | Pin digests; Tempo 3.1 publishes cosign signatures and SLSA provenance |
Grafana Enterprise editions (GEM, GEL, GET) add a gateway with access policies and tokens; open-source users build the equivalent with NGINX, Envoy or an API gateway. Hands-on configuration (TLS, overrides, gateway filters, bucket policies) is in How-to Guides, and the checklist is in Reference.
High Availability and Multi-AZ¶
| Backend | HA mechanism (current majors) |
|---|---|
| Mimir classic | RF=3 across ingesters, zone-aware replication, store-gateway replication |
| Mimir ingest storage | Kafka replication for durability; one ingester per partition per zone for read HA |
| Loki | RF=3 across ingesters (zone-aware), or RF=3 HA monolithic |
| Tempo 3 | Kafka durability (RF1 in Tempo itself); live-stores deployed across zones |
| Pyroscope 2 | Object storage durability; metastore Raft quorum |
Single-AZ deployments are simpler and avoid cross-zone transfer charges, but an AZ outage takes down observability at the moment you need it most. Multi-AZ roughly doubles stateful compute (at least one replica per zone for ingesters, live-stores, store-gateways) while object storage cost stays the same. Kafka adds its own cross-zone replication traffic; Tempo 3.1 supports rack-aware fetching (client_rack, KIP-392) and Mimir has -ingest-storage.kafka.client-rack to keep consumer reads in-zone. A common compromise: multi-AZ for the backends, Grafana on a managed HA database.
Collection and Instrumentation¶
Grafana Alloy is the recommended collector. It embeds upstream OpenTelemetry Collector components (otelcol.*), Prometheus scraping and remote_write, Loki log pipelines (Promtail's code was merged into Alloy), and Pyroscope profiling, configured in the Alloy configuration syntax (formerly River). Grafana Agent reached end of support at the end of 2025, and Promtail was removed from the Loki repository in 3.7.3.
OTel Operator vs Alloy eBPF Auto-Instrumentation¶
| Approach | How it works | Strength | Weakness |
|---|---|---|---|
| OpenTelemetry Operator | Injects language agents into pods via annotations (for example instrumentation.opentelemetry.io/inject-java); manages collectors through CRDs |
Deep, language-level spans and attributes (Java, Node.js, Python, .NET, Go, Apache HTTPD, NGINX) | Per-language agents, pod restarts, language coverage limits |
Alloy beyla.ebpf (Beyla, donated upstream as OpenTelemetry eBPF Instrumentation, OBI) |
eBPF probes observe HTTP/gRPC/SQL traffic of any process on the node | Zero code changes, any language, service-graph breadth | Less application detail; needs privileged access and a recent kernel with BTF |
They complement each other: the Operator for deep traces in key services, eBPF for broad coverage and RED metrics everywhere.
eBPF Maturity for Go¶
Go's register-based ABI (Go 1.17+), static binaries and goroutine scheduling make uprobe-based tracing harder than for C or Rust. Network-level eBPF observability (Cilium/Hubble) and eBPF continuous profiling (Pyroscope's eBPF profiler) are production-proven. eBPF auto-tracing of Go (Beyla/OBI, Odigos) works for HTTP and gRPC boundaries but has gaps around context propagation across goroutines and async patterns, and uprobes on hot paths add overhead. Treat deep function-level eBPF tracing of Go as still maturing (assessment as of 2026-09, not benchmarked here).
Sampling¶
Head sampling (decided when the trace starts, for example parentbased_traceidratio) is cheap but blind to the outcome. Tail sampling waits for the whole trace and can keep every error or slow trace while sampling successes; Alloy provides it through otelcol.processor.tail_sampling, which must see all spans of a trace (use a load-balancing exporter by trace ID in front of the sampling tier) and buffers traces in memory. Tempo 3.x can extrapolate TraceQL metrics from W3C tracestate sampling probability (with(extrapolate=true), 3.1) so sampled data still yields correct rates.
Signal Lifecycle¶
- Instrumentation. The application emits telemetry via OTel SDKs, Prometheus clients or eBPF.
- Collection. Alloy receives, batches, filters, enriches and routes signals.
- Ingestion. Each backend's distributor validates the request, applies tenant limits and shards it (to ingesters, Kafka partitions or segment-writers).
- Durable write. Classic: in-memory + WAL on RF=3 ingesters. Kafka designs: the Kafka log. Pyroscope v2: object storage directly.
- Block building. Ingesters, block-builders or segment-writers cut blocks and upload them (Mimir every 2 h; Tempo and Loki on size/time thresholds).
- Compaction and retention. Compactors (Mimir, Loki), backend workers (Tempo) or compaction-workers (Pyroscope) merge blocks and delete expired data.
- Query. Query-frontends split and cache; queriers read recent data from ingesters/live-stores and history from object storage via store-gateways or directly.
- Visualization. Grafana renders results and cross-signal links.
Why LGTM Is Cheap to Run (and What It Costs Instead)¶
- Object storage is inexpensive. S3 Standard lists at about $0.023/GB-month versus about $0.08/GB-month for EBS gp3 (us-east-1 list prices; verify for your region).
- Label-only indexing (Loki) avoids the storage and CPU of full-text indexes.
- No span index (Tempo); Parquet columns and bloom filters replace an index cluster.
- RF1 with Kafka (Tempo 3, Mimir ingest storage) stores data once instead of three times before compaction.
- No licence fees for the open-source editions.
The bill moves elsewhere: people to operate four or more distributed systems (plus Kafka), cross-AZ network charges, object-storage request charges on cold queries, and cardinality discipline across teams. Rough cost tables are in Reference.
Benchmark Caveats¶
- Published scale figures (for example 1B active series in Mimir) come from Grafana Labs tests in microservices mode on large clusters.
- Loki query speed depends mainly on label selectivity; tens of thousands of active streams per tenant already hurt, and queries without selective labels scan large volumes.
- TraceQL searches over long ranges are slower than indexed trace stores (Jaeger with Elasticsearch) but much cheaper to store.
- Object storage latency varies by provider and region; keep compute in the same region and use caches.
Sources¶
- Mimir architecture: https://grafana.com/docs/mimir/latest/get-started/about-grafana-mimir-architecture/
- Mimir ingest storage: https://grafana.com/docs/mimir/latest/get-started/about-grafana-mimir-architecture/about-ingest-storage-architecture/
- Mimir 3.0 release blog: https://grafana.com/blog/grafana-mimir-3-0-release-all-the-latest-updates/
- Loki deployment modes: https://grafana.com/docs/loki/latest/get-started/deployment-modes/
- Loki overview: https://grafana.com/docs/loki/latest/get-started/overview/
- Tempo architecture: https://grafana.com/docs/tempo/latest/introduction/architecture/
- Tempo 3.0 release blog: https://grafana.com/blog/tempo-3-0-release-all-the-latest-features/
- Migrate from Tempo 2.x to 3.0: https://grafana.com/docs/tempo/latest/set-up-for-tracing/setup-tempo/migrate-to-3/
- Pyroscope v2 architecture: https://grafana.com/docs/pyroscope/latest/reference-pyroscope-v2-architecture/about-pyroscope-v2-architecture/
- Loki multi-tenancy: https://grafana.com/docs/loki/latest/operations/multi-tenancy/
- Tempo multi-tenancy: https://grafana.com/docs/tempo/latest/operations/manage-advanced-systems/multitenancy/
- Mimir authentication and authorization: https://grafana.com/docs/mimir/latest/manage/secure/authentication-and-authorization/