Skip to content

LGTM Stack

Summary

LGTM is Grafana Labs' open-source observability stack: Loki (logs), Grafana (visualization), Tempo (traces) and Mimir (metrics), usually extended with Pyroscope (profiles) and collected by Grafana Alloy. Each backend scales independently, keeps its data in object storage (S3, GCS, Azure Blob), ingests OpenTelemetry natively, and is correlated in Grafana. In 2025-2026 the stack moved to Kafka-decoupled write paths (Mimir 3 ingest storage, Tempo 3), Pyroscope 2 went object-storage-only, the Simple Scalable deployment mode was retired or deprecated, and several Helm charts moved to a community repository. The backends are AGPL-3.0; Alloy is Apache-2.0.

Key Facts

Component Signal Query language Latest Version (date) License Helm chart
Grafana Mimir Metrics PromQL (MQE engine) 3.2.1 (2026-09-10) AGPL-3.0 grafana/mimir-distributed
Grafana Loki Logs LogQL 3.7.8 (2026-09-17) AGPL-3.0 grafana-community/loki
Grafana Tempo Traces TraceQL (+ TraceQL metrics) 3.0.3 (2026-08-13); 3.1.0-rc.1 (2026-09-17) AGPL-3.0 grafana-community/tempo-distributed
Grafana Pyroscope Profiles Label selectors per profile type 2.3.1 (2026-09-08) AGPL-3.0 grafana/pyroscope
Grafana Visualization — 13.2.2 (2026-09-15) AGPL-3.0 grafana-community/grafana
Grafana Alloy Collection Alloy configuration syntax (formerly River) 1.20.0 (2026-09-25) Apache-2.0 grafana/alloy
  • Owner / governance: Grafana Labs (company-led open source, not a CNCF project). Grafana Enterprise Metrics, Logs and Traces (GEM, GEL, GET) are commercial editions built on the same code.
  • Release cadence: Mimir minor releases about quarterly; Loki major about yearly and minor about quarterly; Tempo, Pyroscope and Alloy ship minors every one to two months. Previous minors keep receiving security patches (details in Reference).
  • Recommended install path: Helm on Kubernetes, per component (the lgtm-distributed umbrella chart is deprecated); the grafana/otel-lgtm container for local development.

The compact diagram shows how signals flow; the full architecture is in Explanation.

flowchart LR
    Apps["Apps + infra<br/>(OTel SDKs, exporters,<br/>logs, eBPF)"] --> Alloy["Grafana Alloy"]
    Alloy --> Mimir["Mimir"]
    Alloy --> Loki["Loki"]
    Alloy --> Tempo["Tempo"]
    Alloy --> Pyro["Pyroscope"]
    Mimir --> OS[("Object storage")]
    Loki --> OS
    Tempo --> OS
    Pyro --> OS
    Grafana["Grafana"] -.-> Mimir
    Grafana -.-> Loki
    Grafana -.-> Tempo
    Grafana -.-> Pyro

What Changed in 2025-2026

Date Change Impact
2025-11 Mimir 3.0: ingest storage (Kafka) architecture stable and preferred; Mimir Query Engine default; query-scheduler required; Redis cache and read-write mode removed New Mimir clusters should plan for Kafka; mimir-distributed 6.x deploys ingest storage by default
2025-12 Grafana Agent end of support Use Alloy
2026-01-30 grafana, tempo, tempo-distributed Helm charts moved to grafana-community/helm-charts Change Helm repository URLs
2026-03-16 Loki Helm chart forked to grafana-community/helm-charts Change Helm repository URL; default deploymentMode is now Monolithic
2025-2026 (Loki 3.6-3.7) Loki Simple Scalable Deployment deprecated (removal in Loki 4.0); Promtail removed in Loki 3.7.3 (2026-06-24) Move to HA monolithic or microservices; migrate Promtail to Alloy
2026-04-20 Pyroscope 2.0: v2 storage (segment-writer, metastore, compaction-worker, query-backend) is the default Diskless write path straight to object storage
2026-05-28 Tempo 3.0: ingesters and compactors replaced by block-builders, live-stores and backend scheduler/workers; Kafka required in microservices mode; TraceQL metrics GA Microservices users migrate in parallel; no downgrade

Evaluation

Why it is attractive. It is the most complete open-source stack that covers metrics, logs, traces and profiles with first-class correlation in one UI. Each backend is specialised for its signal, and object storage keeps long retention cheap. Grafana Cloud runs the same code at very large scale, so the components are exercised far beyond typical self-hosted loads.

When it fits:

  • Platform teams with the capacity to operate several distributed systems (now including Kafka for Mimir and Tempo at scale).
  • Organisations standardising on OpenTelemetry and Prometheus that want to avoid SaaS lock-in.
  • Kubernetes environments that need horizontal scale and multi-tenancy for many teams.
  • Cost-sensitive organisations that need long retention.

When it does not fit: small teams without Kubernetes and distributed-systems experience (a single-binary product such as SigNoz or OpenObserve, or a SaaS, is simpler); workloads that need full-text log search without label discipline (Elasticsearch/OpenSearch or VictoriaLogs).

Pros Cons
Each component purpose-built for its signal Operational complexity: 4+ backends, plus Kafka for Mimir/Tempo microservices
Object-storage-first, so long retention is cheap Requires solid Kubernetes and SRE skills
OpenTelemetry-native (OTLP ingestion into every backend) Four query models (PromQL, LogQL, TraceQL, profile selectors)
Large community; proven at Grafana Cloud scale Multi-tenancy requires your own auth gateway in OSS editions
Independent horizontal scaling per component Label cardinality remains the top operational pitfall
Strong cross-signal correlation (exemplars, trace to logs/metrics/profiles) Frequent breaking changes across majors (Tempo 3, Mimir 3, Loki 4 ahead)
grafana/otel-lgtm image for instant local dev Helm chart landscape split across two repositories in 2026

Common use cases:

  • Full-stack Kubernetes observability: metrics, logs, traces and profiles from all workloads.
  • Central multi-tenant observability platform shared by many teams.
  • Cost-effective log aggregation replacing Elasticsearch where queries can start from labels.
  • Distributed tracing at scale with span metrics and service graphs derived from traces.
  • AI/ML pipeline observability: inference latency, GPU utilisation, training metrics.
  • IoT and industrial telemetry: high-volume metric ingestion into Mimir.

Licensing and Pricing

  • Mimir, Loki, Tempo, Pyroscope, Grafana: AGPL-3.0 (Grafana Labs relicensed Grafana, Loki and Tempo from Apache-2.0 in 2021; Mimir launched as AGPL in 2022). Self-hosting is free; if you modify the code and offer it to users over a network, you must publish your modifications under AGPL-3.0.
  • Alloy: Apache-2.0.
  • Grafana Cloud runs the stack as a managed service: Free tier; Pro at a $19/month platform fee plus usage; Enterprise by contract (commonly cited from $25k/year). Prices come from secondary sources as of 2026-09; check https://grafana.com/pricing/. Details and cost models are in Reference and Grafana.

Ecosystem and Data Connections

  • Ingestion protocols: OTLP (gRPC/HTTP) into all backends, Prometheus remote_write (Mimir), Loki push API, Jaeger and Zipkin (Tempo), pprof push (Pyroscope), syslog and Kafka via Alloy.
  • Collection: Grafana Alloy (primary), OpenTelemetry Collector, Prometheus agent mode, Beyla/OpenTelemetry eBPF Instrumentation (OBI). Promtail and Grafana Agent are retired.
  • Storage: S3 and S3-compatible (for example MinIO), GCS, Azure Blob, OpenStack Swift (Mimir); Kafka-compatible logs (Apache Kafka, Redpanda, WarpStream) for Mimir ingest storage and Tempo 3.
  • IaC: Helm charts, Jsonnet/Tanka libraries, Grafana Terraform provider, Tempo Operator and Loki Operator (OpenShift-focused).
  • Instrumentation: OpenTelemetry SDKs and auto-instrumentation agents, Prometheus client libraries, eBPF.

Compatibility and Requirements

  • Kubernetes (recommended; mimir-distributed 6.x requires 1.29+), Docker, or Linux binaries.
  • Object storage is required for Mimir, Loki, Tempo and Pyroscope in production (filesystem only for single-node/dev).
  • Kafka-compatible cluster for Mimir ingest storage (preferred) and Tempo 3 microservices.
  • Memcached for caches; PostgreSQL or MySQL for Grafana's own database.
  • Dev setup: docker run grafana/otel-lgtm (bundles Prometheus, Loki, Tempo, Pyroscope, Grafana and an OTel Collector).

Alternatives

Alternative Model Compared with LGTM
Victoria Stack VictoriaMetrics + VictoriaLogs + VictoriaTraces, Apache-2.0 Lower resource use, local disks instead of object storage, no profiling
SigNoz OTel-native, ClickHouse-backed One storage engine for all signals, fewer moving parts; LGTM scales further per signal and has a larger ecosystem
OpenObserve Rust single binary, object storage Simpler to run; younger ecosystem
Coroot eBPF-first, ClickHouse Opinionated, zero-instrumentation APM
Apache SkyWalking APM with agents Strong Java APM; different data model
ELK / OpenSearch Full-text search Better ad-hoc log search, costlier storage, weaker metrics/traces
Datadog, New Relic, Splunk Observability SaaS Lowest ops burden, highest cost and lock-in

Migration and Lock-in

  • Low lock-in on data paths: OTLP and Prometheus remote_write in, open formats (Prometheus TSDB, Parquet) at rest.
  • Moderate lock-in on queries: PromQL is portable; LogQL and TraceQL are Grafana-specific but well documented.
  • Gradual migration works: dual-ship and move one signal at a time. Mimir accepts remote_write directly; Tempo accepts Jaeger and Zipkin.
  • From ELK: KQL/Lucene queries must be rewritten; Loki's label-only index is a different model.
  • Within LGTM: Tempo 2 to 3 (microservices) is a parallel migration with no downgrade; Loki SSD users must move before 4.0. Recipes are in How-to Guides.

Community Health and Support

  • Very active repositories with frequent releases and security patch lines (see Reference).
  • GitHub stars as of 2026-04: Grafana about 73k, Loki about 28k, Pyroscope about 10k, Mimir about 5k, Tempo about 4k.
  • Public users who presented at Grafana Labs events include Maersk (full LGTM Stack plus Faro for RUM, ObservabilityCON 2023), DHL Express Switzerland (Grafana, Prometheus and k6, Grafana Labs logistics blog) and Salesforce (Grafana, Prometheus and Loki, GrafanaCONline 2021). Checked 2026-09-27; an earlier "Dutch Tax Office" entry was removed because no source was found.
  • Commercial support and SLAs via Grafana Labs (Grafana Cloud, Enterprise editions); community forums and Slack.

Topic Map

  • How-to Guides: deploy, secure, scale, cut cost, migrate, Commands & Recipes, troubleshooting.
  • Reference: versions, Helm charts, deployment modes, components, ports, config keys and defaults, limits, scale figures, hardening checklist.
  • Explanation: design principles, Kafka-decoupled write paths, storage models, correlation, security model, HA trade-offs.

Sources

URL Source kind Authority Retrieved via Date
https://github.com/grafana/mimir repository primary web search 2026-09-25
https://github.com/grafana/mimir/blob/main/CHANGELOG.md changelog primary raw.githubusercontent.com 2026-09-25
https://grafana.com/docs/mimir/latest/ docs primary web search 2026-09-25
https://grafana.com/blog/grafana-mimir-3-0-release-all-the-latest-updates/ release blog primary web search 2026-09-25
https://github.com/grafana/loki repository primary web search 2026-09-25
https://github.com/grafana/loki/releases releases primary web fetch 2026-09-25
https://grafana.com/docs/loki/latest/release-notes/v3-7/ release notes primary raw.githubusercontent.com 2026-09-25
https://grafana.com/docs/loki/latest/get-started/deployment-modes/ docs primary raw.githubusercontent.com 2026-09-25
https://github.com/grafana/tempo repository primary web search 2026-09-25
https://github.com/grafana/tempo/releases releases primary web fetch 2026-09-25
https://grafana.com/blog/tempo-3-0-release-all-the-latest-features/ release blog primary web search 2026-09-25
https://grafana.com/docs/tempo/latest/release-notes/v3-0/ release notes primary web search 2026-09-25
https://grafana.com/docs/tempo/latest/set-up-for-tracing/setup-tempo/migrate-to-3/ docs primary raw.githubusercontent.com 2026-09-25
https://github.com/grafana/pyroscope repository primary web search 2026-09-25
https://grafana.com/docs/pyroscope/latest/release-notes/v2-0/ release notes primary raw.githubusercontent.com 2026-09-25
https://github.com/grafana/alloy repository primary web fetch 2026-09-25
https://grafana.com/docs/alloy/latest/ docs primary web search 2026-04-10
https://github.com/grafana/docker-otel-lgtm repository primary raw.githubusercontent.com 2026-09-25
https://github.com/grafana-community/helm-charts repository primary raw.githubusercontent.com 2026-09-25
https://artifacthub.io/packages/helm/grafana/lgtm-distributed chart registry primary web search 2026-09-25
https://grafana.com/pricing/ pricing primary web search 2026-04-10
https://grafana.com/docs/grafana/latest/datasources/tempo/configure-tempo-data-source/ docs primary web search 2026-04-10
https://grafana.com/docs/grafana/latest/datasources/loki/ docs primary web search 2026-04-10
https://opentelemetry.io/docs/ docs primary web search 2026-04-10
https://signoz.io/comparisons/ comparison secondary web search 2026-04-10
https://www.cloudzero.com/blog/grafana-cloud-pricing/ pricing analysis secondary web search 2026-09-25
https://community.grafana.com/ forum community manual 2026-04-10
https://grafana.com/about/events/grafanacon/ conference primary web search 2026-04-10
https://play.grafana.org/ demo community manual 2026-04-10

Questions

Open

  • When does Loki 3.8 ship, and does it carry the store-backend removals (BoltDB, Cassandra, DynamoDB, BigTable, gRPC) listed on main? Is a Loki 4.0 date announced?
  • Will Loki adopt a Kafka / data-object write path as a supported mode, aligning it with Mimir 3 and Tempo 3?
  • Tempo tenant-ID character and length rules are not documented; confirm they match Mimir's dskit rules (150 bytes).
  • What is the real-world TCO of Kafka (self-managed vs managed vs WarpStream) for Mimir ingest storage and Tempo 3 at 1M series / 50M spans per day?
  • Benchmark object storage request and storage costs across S3, GCS and Azure Blob for LGTM workloads.
  • Deep dive into Grafana Cloud Adaptive Metrics, Logs and Traces savings with measured data.

Resolved

  • All-in-one dev image: grafana/otel-lgtm (OTel Collector, Prometheus, Loki, Tempo, Pyroscope, Grafana). See How-to Guides.
  • Multi-tenancy across LGTM: X-Scope-OrgID on every request, set by an auth gateway. See Explanation.
  • Share one bucket? No; one bucket (or at least prefix) per component. See Reference.
  • Cross-signal correlation: exemplars, derived fields, trace to logs/metrics/profiles, span metrics. See Explanation.
  • Tempo 3.0 architecture: Kafka is required only in microservices mode; monolithic Tempo 3 runs without Kafka (the earlier claim that monolithic also needs Kafka was wrong). See Explanation.
  • Tail sampling in Alloy uses otelcol.processor.tail_sampling, not beyla.ebpf (which is eBPF auto-instrumentation; corrected 2026-09). See How-to Guides.
  • Single-AZ vs multi-AZ: multi-AZ roughly doubles stateful compute; object storage cost unchanged. See Explanation.
  • OTel Operator vs Alloy eBPF auto-instrumentation: complementary depth vs breadth. See Explanation.
  • Limits of monolithic mode at medium scale. See Explanation.
  • SigNoz vs LGTM operational overhead: SigNoz has fewer moving parts; LGTM scales further per signal. See Alternatives.
  • Datadog to LGTM migration gotchas. See How-to Guides.
  • eBPF auto-instrumentation maturity for Go. See Explanation.
  • Loki structured metadata vs labels. See How-to Guides.
  • Grafana Adaptive Metrics/Logs mechanism: usage analysis of queries, dashboards and alerts drives aggregation/drop recommendations. See How-to Guides.