LGTM Stack How-to Guides¶
Task-oriented recipes for running the LGTM stack: deploying it, configuring labels, retention and limits, securing it, scaling it, cutting cost, migrating to it, and troubleshooting it. Version-specific facts (defaults, ports, chart repositories) are tabulated in Reference; background is in Explanation. Grafana server tasks (dashboards, alerting, SSO) belong to Grafana How-to Guides.
Versions assumed
Recipes target Mimir 3.2, Loki 3.7, Tempo 3.0, Pyroscope 2.3, Alloy 1.20 and Grafana 13.2 (September 2026). Several 2026 changes break older guides: Helm chart repositories moved, Tempo 3 removed ingesters and compactors, Promtail was removed, and Mimir 3.0 removed Redis caching.
Deploy a Dev Stack in One Container¶
The grafana/otel-lgtm image bundles an OpenTelemetry Collector, Prometheus (metrics), Tempo, Loki, Pyroscope and Grafana. It is for development, demos and tests only, and it uses Prometheus instead of Mimir.
# Start the stack (Grafana on :3000, OTLP gRPC :4317, OTLP HTTP :4318)
docker run --name lgtm \
-p 3000:3000 \
-p 4317:4317 \
-p 4318:4318 \
--rm -ti grafana/otel-lgtm
# Persist data across restarts
docker run --name lgtm \
-v "$(pwd)/data:/data" \
-p 3000:3000 -p 4317:4317 -p 4318:4318 \
grafana/otel-lgtm
# Show internal component logs (per component: ENABLE_LOGS_LOKI, ENABLE_LOGS_TEMPO, ...)
docker run --name lgtm -e ENABLE_LOGS_ALL=true \
-p 3000:3000 -p 4317:4317 -p 4318:4318 \
grafana/otel-lgtm
- Grafana UI:
http://localhost:3000(default credentialsadmin/admin). - Optional eBPF auto-instrumentation: set
ENABLE_OBI=true(needs Linux 5.8+ with BTF,--pid=host --privileged; the repo'srun-lgtm.shadds the flags).
Send a Test Trace¶
curl -X POST http://localhost:4318/v1/traces \
-H "Content-Type: application/json" \
-d '{
"resourceSpans": [{
"resource": {"attributes": [{"key": "service.name", "value": {"stringValue": "test-service"}}]},
"scopeSpans": [{
"spans": [{
"traceId": "5b8efff798038103d269b633813fc60c",
"spanId": "eee19b7ec3c1b174",
"name": "test-span",
"kind": 1,
"startTimeUnixNano": "1544712660000000000",
"endTimeUnixNano": "1544712661000000000"
}]
}]
}]
}'
Open Grafana, go to Explore, choose the Tempo data source and search for trace ID 5b8efff798038103d269b633813fc60c.
Deploy to Kubernetes with Helm¶
Charts now come from two repositories (see Reference): grafana keeps mimir-distributed, pyroscope, alloy and k8s-monitoring; grafana-community now hosts loki, tempo, tempo-distributed and grafana. Do not use the deprecated lgtm-distributed umbrella chart.
# Add both repositories
helm repo add grafana https://grafana.github.io/helm-charts
helm repo add grafana-community https://grafana-community.github.io/helm-charts
helm repo update
# Deploy storage backends first, then collectors, then Grafana
helm install mimir grafana/mimir-distributed -n monitoring --create-namespace -f mimir-values.yaml
helm install loki grafana-community/loki -n monitoring -f loki-values.yaml
helm install tempo grafana-community/tempo-distributed -n monitoring -f tempo-values.yaml
helm install pyroscope grafana/pyroscope -n monitoring -f pyroscope-values.yaml
helm install alloy grafana/alloy -n monitoring -f alloy-values.yaml
helm install grafana grafana-community/grafana -n monitoring -f grafana-values.yaml
# Community charts are also available as OCI artifacts
helm install loki oci://ghcr.io/grafana-community/helm-charts/loki -n monitoring -f loki-values.yaml
Existing releases on the old repository
Releases installed from grafana/loki, grafana/tempo-distributed or grafana/grafana keep working, but the old repository no longer receives new versions or security fixes for those charts. Point helm upgrade at grafana-community/<chart> with the same release name and values, after reading the chart's upgrade notes.
Production Readiness Checklist¶
- Object storage configured for Mimir, Loki, Tempo and Pyroscope (separate buckets)
- Kafka-compatible cluster sized for Mimir ingest storage and Tempo 3 microservices (or Mimir classic architecture chosen deliberately)
- PostgreSQL or MySQL for Grafana metadata (not SQLite)
- Memcached for results, chunks, index and metadata caches
- Auth gateway that sets
X-Scope-OrgIDfrom verified identity - Ingress / load balancer with TLS termination
- Autoscaling (HPA or KEDA) and resource requests/limits on every component
- Zone-aware replication for stateful components (ingesters, live-stores, store-gateways)
- Provisioning for data sources, dashboards and alert rules in Git
- Cross-signal correlation configured (exemplars, trace to logs, derived fields, trace to profiles)
- Retention set per backend (the defaults keep Mimir and Loki data forever)
- Meta-monitoring stack in place for the LGTM components themselves
Instrument Applications for Correlation¶
- Standardise on OpenTelemetry and OTLP for new services; keep Prometheus client libraries where they already exist.
- Inject trace context into logs (OTel log bridges or logging-library integrations) so every log line carries
trace_idandspan_id. - Set resource attributes on every signal:
service.name,service.namespace,deployment.environment.name,k8s.namespace.name,k8s.pod.name. Consistentservice.nameis what makes trace-to-logs and trace-to-profiles queries match. - Start with auto-instrumentation (Java agent, Python
opentelemetry-instrument, .NET and Node.js agents, eBPF via Alloybeyla.ebpffor anything else), then add manual spans for business-critical logic. - Enable span profiling (Pyroscope SDK span-profile integrations) for services where CPU hotspots matter.
- Sample deliberately in production (see Sample Traces).
- Govern labels and GitOps everything: enforce naming and cardinality rules in review, and keep dashboards, alert rules, data sources and Helm values in version control.
Migrate Loki from Simple Scalable to HA Monolithic or Microservices¶
Loki's Simple Scalable Deployment (read/write/backend) is deprecated and will not run in Loki 4.0.
- Decide the target: HA monolithic for modest volume (tens of GB/day) with HA, or Distributed (microservices) for large volume.
- For HA monolithic, run at least three
-target=allreplicas withcommon.replication_factor: 3, memberlist ring (common.ring.kvstore.store: memberlist), shared object storage, and a load balancer in front. Setcompactor.horizontal_scaling_mode: mainon one instance,workeron the rest, and pointcommon.compactor_grpc_addressat the main one (see the upstreamexamples/ha-monolithic). - For Helm, change
deploymentModestep by step using the chart's transitional modes (SimpleScalable<->Distributed), wait until the new components are ready, then switch toDistributedand scaleread,writeandbackendto zero. - Keep the same schema config and buckets; data in object storage does not need migration.
Configure Labels and Structured Metadata¶
Label cardinality is the most common Loki (and Mimir) failure mode. Use labels for how you select data, structured metadata for what you filter within it.
| Put in | Good examples | Bad examples |
|---|---|---|
| Stream labels (indexed, low cardinality, about 10 or fewer per stream) | namespace, cluster, service_name, env, job, container |
user_id, request_id, trace_id, pod UID, ip_address, url_path |
| Structured metadata (not indexed, per line) | trace_id, span_id, k8s.pod.name, service.instance.id, user_id |
Large payloads (stack traces: drop or keep in the line) |
| Log line | The message and its fields | — |
Steps:
- Use Loki schema
v13with the TSDB index;allow_structured_metadatadefaults totruethere. - Ship logs over OTLP (Alloy
otelcol.exporter.otlphttptohttp://<loki>/otlp) so resource attributes become labels or structured metadata automatically. - Tune the mapping with
limits_config.otlp_config(promote a few resource attributes to labels, drop noisy attributes withaction: drop). - Never promote
k8s.pod.nameorservice.instance.idto an index label. - Query structured metadata directly, without parsers:
{service_name="checkout"} | trace_id="abc123". - Watch active streams per tenant (
max_global_streams_per_user, default 5000); keep them in the low tens of thousands at most.
Configure Retention¶
Retention is off by default for Mimir and Loki. Set it explicitly (keys and defaults in Reference).
# Mimir: keep blocks 13 months (per-tenant overrides possible)
limits:
compactor_blocks_retention_period: 395d # about 13 months; note: not blocks_storage.tsdb.retention_period (ingester-local, 13h)
# Loki: 30 days; retention requires the compactor to enforce it
limits_config:
retention_period: 720h
compactor:
retention_enabled: true
delete_request_store: s3
# Tempo 3.x: 14 days (backend workers enforce it)
backend_worker:
compaction:
block_retention: 336h
# Pyroscope 2.x (v2 storage)
limits:
retention_period: 14d
Pair retention with bucket lifecycle rules only for data the backends no longer manage (for example Tempo's compacted-block leftovers); never let a bucket lifecycle rule delete blocks a backend still tracks.
Set Per-Tenant Limits¶
Per-tenant overrides live in a runtime file that each backend reloads without a restart (Mimir runtime_config.file, Loki runtime_config.file, Tempo overrides.per_tenant_override_config).
# Mimir runtime config
overrides:
tenant-engineering:
ingestion_rate: 100000 # samples/s
ingestion_burst_size: 200000
max_global_series_per_user: 500000
max_fetched_chunks_per_query: 2000000
max_global_exemplars_per_user: 100000
compactor_blocks_retention_period: 400d
# Loki runtime config
overrides:
tenant-frontend:
ingestion_rate_mb: 10
max_global_streams_per_user: 20000
max_chunks_per_query: 100000
max_query_length: 721h
retention_period: 720h
# Tempo 3.x per-tenant overrides (scoped format; the legacy flat format is disabled by default in 3.0)
overrides:
"tenant-payments":
ingestion:
rate_limit_bytes: 50000000
burst_size_bytes: 100000000
max_traces_per_user: 20000
global:
max_bytes_per_trace: 5000000
metrics_generator:
processors: [span-metrics, service-graphs]
Tempo ships tempo-cli migrate overrides-config to convert legacy overrides files.
Security Hardening¶
Put an Auth Gateway in Front¶
- Deploy NGINX, Envoy or an API gateway (Grafana Enterprise editions include one) in front of every distributor and query-frontend.
- Authenticate clients (tokens, mTLS or SSO), strip any client-supplied
X-Scope-OrgID, and set it from the verified identity. - Expose only the gateway; keep all backend services cluster-internal.
Example Envoy external-authorization filter (the auth service decides and returns the tenant header):
http_filters:
- name: envoy.filters.http.ext_authz
typed_config:
"@type": type.googleapis.com/envoy.extensions.filters.http.ext_authz.v3.ExtAuthz
http_service:
server_uri:
uri: auth-service:8080
cluster: ext-authz
timeout: 0.25s
authorization_response:
allowed_upstream_headers:
patterns:
- exact: x-scope-orgid
Send Tenant Headers from Collectors¶
# OpenTelemetry Collector: OTLP export to Tempo with a tenant header
exporters:
otlp:
endpoint: tempo-distributor.monitoring.svc:4317
headers:
X-Scope-OrgID: tenant-engineering
// Grafana Alloy: remote_write to Mimir with a tenant header
prometheus.remote_write "mimir" {
endpoint {
url = "https://lgtm-gateway.example.com/api/v1/push"
headers = { "X-Scope-OrgID" = "tenant-engineering" }
}
}
Enable TLS¶
Mimir supports TLS (and client-certificate auth) on both HTTP and gRPC:
-server.http-tls-cert-path=/certs/server.crt
-server.http-tls-key-path=/certs/server.key
-server.http-tls-client-auth="RequireAndVerifyClientCert"
-server.http-tls-ca-path="/certs/ca.crt"
-server.grpc-tls-cert-path=/certs/server.crt
-server.grpc-tls-key-path=/certs/server.key
-server.grpc-tls-client-auth="RequireAndVerifyClientCert"
-server.grpc-tls-ca-path="/certs/ca.crt"
Tempo uses the same dskit server block in YAML:
server:
http_listen_port: 3200
grpc_listen_port: 9095
http_tls_config:
cert_file: /etc/tempo/certs/server.crt
key_file: /etc/tempo/certs/server.key
client_auth_type: RequireAndVerifyClientCert
client_ca_file: /etc/tempo/certs/ca.crt
grpc_tls_config:
cert_file: /etc/tempo/certs/server.crt
key_file: /etc/tempo/certs/server.key
client_auth_type: RequireAndVerifyClientCert
client_ca_file: /etc/tempo/certs/ca.crt
The OTLP receivers on Tempo distributors take their own tls settings under distributor.receivers.otlp.protocols.<grpc|http>.tls. For Kafka, enable SASL and TLS on both Mimir (-ingest-storage.kafka.sasl-*, -ingest-storage.kafka.tls-enabled) and Tempo (ingest.kafka.sasl_mechanism, tls_enabled, Tempo 3.1+).
Encrypt and Isolate Object Storage¶
Choose server-side encryption per bucket:
| Method | Description | Key management |
|---|---|---|
| SSE-S3 | AWS-managed keys | AWS handles rotation |
| SSE-KMS | Customer-managed KMS key | Full key lifecycle, CloudTrail audit |
| SSE-C | Customer-provided key | Client-side key management |
| CSE | Client-side encryption | Application encrypts before upload |
Recommendation
Use SSE-KMS with a customer-managed key per component. You get an audit trail of key usage and can rotate Mimir, Loki, Tempo and Pyroscope keys independently.
Use one bucket per component per environment and restrict each to that component's role:
observability-mimir-<env> # metrics blocks (plus -ruler, -alertmanager)
observability-loki-<env> # log chunks and index
observability-tempo-<env> # trace blocks
observability-pyroscope-<env> # profile segments and blocks
{
"Effect": "Allow",
"Principal": {"AWS": "arn:aws:iam::123456789012:role/mimir-role"},
"Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject", "s3:ListBucket"],
"Resource": ["arn:aws:s3:::observability-mimir-prod", "arn:aws:s3:::observability-mimir-prod/*"]
}
On EKS, bind these roles to service accounts (IRSA or EKS Pod Identity) instead of static access keys.
Configure Tempo Operator OIDC Multi-Tenancy¶
The Tempo Operator (grafana/tempo-operator, used mainly on OpenShift) can front Tempo with a gateway that does OIDC authentication and static RBAC. Check that your operator release supports Tempo 3.x before relying on it.
apiVersion: tempo.grafana.com/v1alpha1
kind: TempoStack
metadata:
name: production
spec:
template:
gateway:
enabled: true
queryFrontend:
jaegerQuery:
enabled: true
tenants:
mode: static
authentication:
- tenantName: engineering
tenantId: eng-team
oidc:
issuerURL: https://dex.example.com/dex
redirectURL: https://tempo-gateway.example.com/oidc/eng-team/callback
usernameClaim: email
secret:
name: tempo-oidc-secret
authorization:
roles:
- name: read-write
permissions: [read, write]
resources: [traces]
tenants: [engineering]
roleBindings:
- name: engineers
roles: [read-write]
subjects:
- kind: user
name: eng@example.com
Configure Grafana Data Sources for Tenants¶
- Set
X-Scope-OrgIDas a custom HTTP header on each data source, or create one data source per tenant. - If the gateway requires credentials, enable basic auth or forward the user's OAuth token on the data source.
- For Loki cross-tenant queries, enable
querier.multi_tenant_queries_enabled: true, sendX-Scope-OrgID: a|b, and filter with the__tenant_id__label ({app="api", __tenant_id__=~"eng.+"}). Push (POST /loki/api/v1/push) and tail (GET /loki/api/v1/tail) reject multi-tenant requests.
Restrict Network Paths¶
- Kubernetes NetworkPolicies: only collectors (Alloy) reach distributors; only Grafana and the gateway reach query-frontends; only Mimir/Tempo pods reach Kafka.
- Reach object storage through VPC endpoints or private links, never the public internet.
- Do not expose Kafka, Memcached or memberlist ports outside the cluster.
Scale and Operate¶
Scaling Decision Matrix¶
| Symptom | Component to scale | How |
|---|---|---|
| Slow metric queries | Mimir queriers / store-gateways | Add replicas; check query sharding (default on in 3.2) and caches |
| Growing Kafka consumer lag (Mimir) | Mimir ingesters (partitions) | Add partitions and ingesters per zone; watch cortex_ingest_storage_reader_receive_and_consume_delay_seconds |
| Metric write errors (classic) | Mimir distributors / ingesters | Add replicas; check per-tenant limits |
| Slow log search | Loki queriers | Add replicas; tighten label selectors; check stream cardinality |
| Log ingestion lag or 429s | Loki distributors / ingesters | Add replicas; raise ingestion_rate_mb only after checking cardinality |
| Recent trace queries failing (Tempo 3) | Live-stores | Add partitions + live-stores (replicas must equal partitions); KEDA scaling on bytes held |
| Slow TraceQL search | Tempo queriers | Add replicas; narrow time range; add dedicated columns |
| Cache miss rate above 20% | Memcached | Add replicas or memory |
| Object storage latency | All | Same-region compute, caches, VPC endpoints |
High Availability Minimums¶
| Component | HA mechanism | Minimum |
|---|---|---|
| Mimir ingesters (classic) / Loki ingesters | RF=3, zone-aware | 3 (one per zone) |
| Mimir ingesters (ingest storage) | One per partition per zone | 2 zones |
| Tempo live-stores | Zone-aware, replicas = partitions | 2 zones |
| Kafka | Broker replication | 3 brokers across zones |
| Distributors, queriers, query-frontends, query-schedulers | Stateless, load balanced | 2+ |
| Mimir store-gateways | Sharded + replicated | 3 |
| Compactors (Mimir, Loki) / Tempo backend-scheduler | Sharded or singleton | 1+ (Tempo scheduler is a singleton) |
| Pyroscope metastore | Raft | 3 |
Monitor the Monitoring¶
- Run a separate, small meta-monitoring stack (or Grafana Cloud's free tier) that watches the primary stack, so an outage does not blind you.
- Install the mixins (dashboards, alerts, runbooks):
operations/mimir-mixin/ingrafana/mimir,production/loki-mixin/ingrafana/loki,operations/tempo-mixin/ingrafana/tempo. - Collect component metrics and logs with the
k8s-monitoringchart (manages Alloy); themimir-distributedchart deprecated its built-in Grafana Agent meta-monitoring in 6.0. - Alert on Kafka lag, ingestion errors, compaction failures and object storage errors; key metric names are in Reference.
- Load-test before production with
k6and the Mimir/Loki/Tempo test tooling (mimir-continuous-test,loki-canary,tempo-vulture).
Reduce Cost¶
| Cost driver | Primary factor | Optimization |
|---|---|---|
| Object storage | Volume x retention | Retention per tenant, compression, avoid tiny objects |
| Compute (write path) | Ingestion rate, cardinality | Drop unused series/labels in Alloy; right-size; spot nodes for stateless parts |
| Compute (read path) | Query volume and shape | Recording rules, caches, query limits |
| Kafka | Throughput x replication | Compression (-ingest-storage.kafka.producer-compression), rack-aware fetching, managed or diskless Kafka (for example WarpStream) |
| Network (cross-AZ) | Replication and queries across zones | Zone-aware routing, client_rack / -ingest-storage.kafka.client-rack, VPC endpoints |
| Memcached | Cache size vs hit ratio | Size for above 80% hit rate |
Strategies, roughly in order of payoff:
- Filter at the edge. Drop debug logs, unused metrics and high-cardinality labels in Alloy before they reach the backends.
- Recording rules for expensive dashboards and alerts in Mimir.
- Trace sampling (next section) to cut Tempo volume.
- Per-tenant retention instead of one global long retention.
- Grafana Cloud Adaptive Metrics / Adaptive Logs / Adaptive Traces if you use Grafana Cloud: they analyse which series and logs are queried or alerted on and recommend aggregation or drop rules (vendor claim: at launch Grafana Labs reported that tests in more than 150 customer environments cut time-series volume by 20-50% on average, per the Adaptive Metrics announcement; no independent measurement is published).
- Single-AZ for non-critical environments.
- Storage lifecycle tiers only for data the backend no longer reads.
Sample Traces¶
- Head sampling in SDKs:
OTEL_TRACES_SAMPLER=parentbased_traceidratio,OTEL_TRACES_SAMPLER_ARG=0.1. - Tail sampling in Alloy keeps every error and slow trace. All spans of a trace must reach the same Alloy instance, so put
otelcol.exporter.loadbalancing(routing by trace ID) in front of the sampling tier.
otelcol.processor.tail_sampling "default" {
decision_wait = "10s"
policy {
name = "keep-errors"
type = "status_code"
status_code {
status_codes = ["ERROR"]
}
}
policy {
name = "keep-slow"
type = "latency"
latency {
threshold_ms = 2000
}
}
policy {
name = "sample-the-rest"
type = "probabilistic"
probabilistic {
sampling_percentage = 10
}
}
output {
traces = [otelcol.exporter.otlp.tempo.input]
}
}
Tail sampling buffers traces in memory for decision_wait, so size Alloy memory accordingly. If SDKs record the sampling probability in W3C tracestate, Tempo 3.1 TraceQL metrics can extrapolate true rates with with(extrapolate=true).
Migrate to or Within LGTM¶
Prometheus to Mimir¶
# Add to the existing Prometheus config: zero-downtime, dual-write while you validate
remote_write:
- url: http://mimir-distributor:8080/api/v1/push
headers:
X-Scope-OrgID: default
Backfill historical TSDB blocks with mimirtool backfill if you need old data.
Jaeger to Tempo¶
Tempo distributors accept Jaeger (Thrift HTTP on 14268, gRPC on 14250 when the Jaeger receiver is enabled), Zipkin (9411) and OTLP (4317/4318). Re-point Jaeger clients or collectors, then move them to OTLP; Jaeger client libraries are retired upstream.
Promtail to Alloy¶
Promtail was removed from the Loki repository in 3.7.3. Convert its configuration:
alloy convert --source-format=promtail --output=config.alloy promtail.yaml
# or run a Promtail config directly while migrating
alloy run --config.format=promtail promtail.yaml
Tempo 2.x to 3.0¶
- Monolithic: convert the config and upgrade the binary; no Kafka needed.
- Microservices: provision Kafka, deploy 3.0 in parallel on the same bucket, switch write and read traffic, then decommission 2.x. There is no in-place downgrade.
- Before upgrading: change any
vParquet3write format tovParquet4or later; migrate legacy overrides (tempo-cli migrate overrides-config); move off the removedscalable-single-binarytarget; removequerier.query_live_storeandquery_frontend.search.query_ingesters_until.
Elasticsearch/Kibana to Loki/Grafana¶
- Deploy Loki alongside Elasticsearch.
- Have Alloy send logs to both (dual-write).
- Rebuild critical Kibana dashboards in Grafana with LogQL (KQL/Lucene queries must be rewritten; Loki has no full-text index).
- Validate completeness and query parity.
- Stop writing to Elasticsearch and decommission it after its retention expires.
Datadog to Self-Hosted LGTM¶
- Agent: replace the Datadog Agent with Alloy. DogStatsD extensions do not map one-to-one, and Datadog's auto-discovered integrations must be replaced with explicit Prometheus exporters/scrape configs.
- Tracing libraries: move from
dd-trace-*to OpenTelemetry SDKs; Datadog's proprietary propagation must be bridged to W3C TraceContext during the transition. - Tags to labels: sanitize names (Prometheus labels allow no dots) and prune high-cardinality tags.
- Dashboards: no automatic converter; rebuild in Grafana.
- Alerts: rewrite monitors in PromQL/LogQL and map notification targets to Grafana contact points.
- Order: dual-ship, then migrate infrastructure metrics, application metrics, logs, traces, alerts and dashboards; cut over per team.
Commands & Recipes¶
Grafana Data Source Provisioning (Cross-Signal)¶
This provisioning file wires every cross-signal link described in Explanation. URLs assume in-cluster service names; add X-Scope-OrgID headers when multi-tenancy is on.
# /etc/grafana/provisioning/datasources/lgtm.yaml
apiVersion: 1
datasources:
# === METRICS (Mimir) ===
- name: Mimir
type: prometheus
uid: mimir
access: proxy
url: http://mimir-query-frontend:8080/prometheus
isDefault: true
jsonData:
httpMethod: POST
exemplarTraceIdDestinations:
- name: traceID
datasourceUid: tempo
# === LOGS (Loki) ===
- name: Loki
type: loki
uid: loki
access: proxy
url: http://loki-gateway
jsonData:
derivedFields:
- datasourceUid: tempo
matcherType: label # read trace_id from structured metadata
matcherRegex: trace_id
name: TraceID
url: '$${__value.raw}'
urlDisplayLabel: 'View Trace'
# === TRACES (Tempo) ===
- name: Tempo
type: tempo
uid: tempo
access: proxy
url: http://tempo-query-frontend:3200
jsonData:
tracesToLogsV2:
datasourceUid: loki
spanStartTimeShift: '-1h'
spanEndTimeShift: '1h'
tags:
- key: service.name
value: service_name
filterByTraceID: true
filterBySpanID: false
tracesToMetrics:
datasourceUid: mimir
spanStartTimeShift: '-1h'
spanEndTimeShift: '1h'
tags:
- key: service.name
value: service
queries:
- name: 'Request Rate'
query: 'sum(rate(traces_spanmetrics_calls_total{$$__tags}[5m]))'
- name: 'Error Rate'
query: 'sum(rate(traces_spanmetrics_calls_total{$$__tags,status_code="STATUS_CODE_ERROR"}[5m]))'
tracesToProfiles:
datasourceUid: pyroscope
tags:
- key: service.name
value: service_name
profileTypeId: 'process_cpu:cpu:nanoseconds:cpu:nanoseconds'
serviceMap:
datasourceUid: mimir
nodeGraph:
enabled: true
# === PROFILES (Pyroscope) ===
- name: Pyroscope
type: grafana-pyroscope-datasource
uid: pyroscope
access: proxy
url: http://pyroscope:4040
If trace IDs are embedded in the log line instead of structured metadata, use a regex matcher: matcherType: regex with matcherRegex: '"traceID":"(\w+)"'.
Alloy Configuration (Full LGTM Pipeline)¶
Receives OTLP and scrapes Kubernetes pods, then routes metrics to Mimir, logs to Loki's native OTLP endpoint and traces to Tempo. memory_limiter comes first, as upstream recommends.
// config.alloy: metrics, logs and traces to Mimir, Loki and Tempo
// ---------- Receivers ----------
otelcol.receiver.otlp "default" {
grpc { endpoint = "0.0.0.0:4317" }
http { endpoint = "0.0.0.0:4318" }
output {
metrics = [otelcol.processor.memory_limiter.default.input]
logs = [otelcol.processor.memory_limiter.default.input]
traces = [otelcol.processor.memory_limiter.default.input]
}
}
discovery.kubernetes "pods" {
role = "pod"
}
prometheus.scrape "k8s_pods" {
targets = discovery.kubernetes.pods.targets
forward_to = [prometheus.remote_write.mimir.receiver]
}
// ---------- Processors ----------
otelcol.processor.memory_limiter "default" {
check_interval = "1s"
limit = "512MiB"
output {
metrics = [otelcol.processor.batch.default.input]
logs = [otelcol.processor.batch.default.input]
traces = [otelcol.processor.batch.default.input]
}
}
otelcol.processor.batch "default" {
output {
metrics = [otelcol.exporter.prometheus.mimir.input]
logs = [otelcol.exporter.otlphttp.loki.input]
traces = [otelcol.exporter.otlp.tempo.input]
}
}
// ---------- Exporters ----------
otelcol.exporter.prometheus "mimir" {
forward_to = [prometheus.remote_write.mimir.receiver]
}
prometheus.remote_write "mimir" {
endpoint {
url = "http://mimir-distributor:8080/api/v1/push"
headers = { "X-Scope-OrgID" = "default" }
}
}
// Loki native OTLP ingestion (resource attributes -> labels / structured metadata)
otelcol.exporter.otlphttp "loki" {
client {
endpoint = "http://loki-gateway/otlp"
headers = { "X-Scope-OrgID" = "default" }
}
}
otelcol.exporter.otlp "tempo" {
client {
endpoint = "tempo-distributor:4317"
tls { insecure = true }
}
}
Helm Values Snippets¶
Mimir (grafana/mimir-distributed 6.x, ingest storage)¶
# mimir-values.yaml (key settings only)
kafka:
enabled: false # the bundled single-node Kafka is for demos only
minio:
enabled: false # use real object storage
mimir:
structuredConfig:
common:
storage:
backend: s3
s3:
endpoint: s3.us-east-1.amazonaws.com
region: us-east-1
blocks_storage:
s3:
bucket_name: observability-mimir-blocks
ruler_storage:
s3:
bucket_name: observability-mimir-ruler
alertmanager_storage:
s3:
bucket_name: observability-mimir-alertmanager
ingest_storage:
kafka:
address: kafka-bootstrap.kafka.svc:9092
topic: mimir-ingest
limits:
max_global_series_per_user: 1500000
ingestion_rate: 100000
compactor_blocks_retention_period: 395d
ingester:
zoneAwareReplication:
enabled: true
resources:
requests: { cpu: "1", memory: "4Gi" }
limits: { memory: "8Gi" }
persistentVolume:
enabled: true
size: 50Gi
querier:
replicas: 2
store_gateway:
zoneAwareReplication:
enabled: true
compactor:
replicas: 1
To stay on the classic architecture, set mimir.structuredConfig.ingest_storage.enabled: false and follow the chart's migration guide; switching architectures on a live cluster needs a planned migration.
Loki (grafana-community/loki, Distributed mode)¶
# loki-values.yaml
deploymentMode: Distributed
loki:
auth_enabled: true
schemaConfig:
configs:
- from: "2024-04-01"
store: tsdb
object_store: s3
schema: v13
index:
prefix: loki_index_
period: 24h
storage:
type: s3
bucketNames:
chunks: observability-loki-chunks
ruler: observability-loki-ruler
s3:
region: us-east-1
limits_config:
retention_period: 720h # 30 days
max_global_streams_per_user: 10000
ingestion_rate_mb: 20
per_stream_rate_limit: 5MB
allow_structured_metadata: true
compactor:
retention_enabled: true
delete_request_store: s3
ingester:
replicas: 3
distributor:
replicas: 2
querier:
replicas: 2
queryFrontend:
replicas: 2
queryScheduler:
replicas: 2
indexGateway:
replicas: 2
compactor:
replicas: 1
# Disable the other modes' workloads
singleBinary:
replicas: 0
read:
replicas: 0
write:
replicas: 0
backend:
replicas: 0
Tempo (grafana-community/tempo-distributed 3.x, Tempo 3)¶
# tempo-values.yaml
multitenancyEnabled: true
storage:
trace:
backend: s3
s3:
bucket: observability-tempo-traces
endpoint: s3.us-east-1.amazonaws.com
region: us-east-1
ingest:
kafka:
address: kafka-bootstrap.kafka.svc:9092
topic: tempo-traces
auto_create_topic_default_partitions: 3
blockBuilder:
replicas: 3 # must equal the partition count
liveStore:
replicas: 3 # must equal the partition count
metricsGenerator:
enabled: true
config:
storage:
remote_write:
- url: http://mimir-distributor:8080/api/v1/push
send_exemplars: true
overrides:
defaults:
metrics_generator:
processors: [span-metrics, service-graphs]
querier:
replicas: 2
OpenTelemetry SDK Quickstart¶
Java (Auto-Instrumentation)¶
# Download the OpenTelemetry Java agent
curl -LO https://github.com/open-telemetry/opentelemetry-java-instrumentation/releases/latest/download/opentelemetry-javaagent.jar
# Run the app with the agent
java -javaagent:opentelemetry-javaagent.jar \
-Dotel.service.name=my-service \
-Dotel.exporter.otlp.endpoint=http://alloy:4318 \
-jar my-app.jar
The Java agent defaults to OTLP over HTTP/protobuf, hence port 4318.
Python (Auto-Instrumentation)¶
pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install
OTEL_SERVICE_NAME=my-service \
OTEL_EXPORTER_OTLP_ENDPOINT=http://alloy:4317 \
OTEL_EXPORTER_OTLP_PROTOCOL=grpc \
opentelemetry-instrument python app.py
Go (Manual SDK)¶
import (
"context"
"go.opentelemetry.io/otel"
"go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc"
"go.opentelemetry.io/otel/sdk/resource"
sdktrace "go.opentelemetry.io/otel/sdk/trace"
semconv "go.opentelemetry.io/otel/semconv/v1.26.0"
)
func initTracer(ctx context.Context) (*sdktrace.TracerProvider, error) {
exporter, err := otlptracegrpc.New(ctx,
otlptracegrpc.WithEndpoint("alloy:4317"),
otlptracegrpc.WithInsecure(),
)
if err != nil {
return nil, err
}
tp := sdktrace.NewTracerProvider(
sdktrace.WithBatcher(exporter),
sdktrace.WithResource(resource.NewWithAttributes(
semconv.SchemaURL,
semconv.ServiceName("my-service"),
)),
)
otel.SetTracerProvider(tp)
return tp, nil
}
Environment Variables (All Languages)¶
export OTEL_SERVICE_NAME=my-service
export OTEL_EXPORTER_OTLP_ENDPOINT=http://alloy:4317
export OTEL_EXPORTER_OTLP_PROTOCOL=grpc
export OTEL_RESOURCE_ATTRIBUTES="deployment.environment.name=production,k8s.namespace.name=default"
export OTEL_TRACES_SAMPLER=parentbased_traceidratio
export OTEL_TRACES_SAMPLER_ARG=0.1 # 10% head sampling
Useful One-Liners¶
# Readiness of each backend (ports differ per component)
curl -s http://mimir-distributor:8080/ready
curl -s http://loki-distributor:3100/ready
curl -s http://tempo-distributor:3200/ready
# Query Mimir directly
curl -s -H "X-Scope-OrgID: default" \
"http://mimir-query-frontend:8080/prometheus/api/v1/query?query=up" | jq .
# Push a test log to Loki
curl -X POST -H "Content-Type: application/json" \
-H "X-Scope-OrgID: default" \
"http://loki-distributor:3100/loki/api/v1/push" \
-d '{"streams":[{"stream":{"app":"test"},"values":[["'$(date +%s)000000000'","hello from curl"]]}]}'
# Query Loki directly
curl -s -G -H "X-Scope-OrgID: default" \
"http://loki-query-frontend:3100/loki/api/v1/query_range" \
--data-urlencode 'query={app="test"}' --data-urlencode 'limit=10' | jq .
# Fetch a trace from Tempo by ID
curl -s "http://tempo-query-frontend:3200/api/v2/traces/5b8efff798038103d269b633813fc60c" | jq .
# Mimir: show runtime-config overrides that differ from defaults
curl -s "http://mimir-distributor:8080/runtime_config?mode=diff"
# Validate an Alloy configuration before rollout
alloy fmt config.alloy
Troubleshooting¶
| Symptom | Likely cause | Fix |
|---|---|---|
429 / "ingestion rate limit exceeded" |
Tenant over ingestion_rate / ingestion_rate_mb |
Raise the per-tenant limit or reduce volume at the collector |
| "max streams limit reached" (Loki) | High label cardinality | Move high-cardinality values to structured metadata; drop labels in Alloy |
| "per-user series limit" (Mimir) | Too many active series | Drop or aggregate series; raise max_global_series_per_user deliberately |
| "context deadline exceeded" on queries | Slow object storage or oversized query | Enable caches, add query limits, check region placement |
| Exemplars not showing | Exemplar storage disabled in Mimir | Set max_global_exemplars_per_user > 0; confirm the app emits exemplars and Prometheus data source has exemplarTraceIdDestinations |
| Trace to logs not working | No trace ID in logs | Make the OTel SDK/log bridge inject trace_id; ship logs via OTLP so it lands in structured metadata |
| Derived fields not clickable | Matcher does not match | Use matcherType: label for structured metadata; test regexes against real lines |
| Recent traces missing or query errors (Tempo 3) | Live-store lagging behind Kafka | Scale live-stores; check Kafka health; query_frontend.query_end_cutoff (30 s) hides the newest data by design |
| Tempo 3 refuses to start with "legacy overrides" error | Flat overrides format | Run tempo-cli migrate overrides-config, or temporarily enable_legacy_overrides: true |
| High memory on ingesters / live-stores | Too many series, streams or live traces | Scale out; tune limits; check for cardinality spikes |
| Growing Kafka lag (Mimir ingest storage) | Too few partitions/ingesters or slow disks | Add partitions and ingesters per zone; check ingester CPU and disk |
| Slow TraceQL queries | Wide time range, low selectivity | Narrow range, filter on dedicated columns, use TraceQL metrics instead of search |
Sources¶
- Grafana community Helm charts: https://github.com/grafana-community/helm-charts
mimir-distributedchart docs: https://grafana.com/docs/helm-charts/mimir-distributed/latest/- Loki Helm install docs: https://grafana.com/docs/loki/latest/setup/install/helm/
- Loki OTLP ingestion: https://grafana.com/docs/loki/latest/send-data/otel/
- Loki retention: https://grafana.com/docs/loki/latest/operations/storage/retention/
- Tempo 2.x to 3.0 migration: https://grafana.com/docs/tempo/latest/set-up-for-tracing/setup-tempo/migrate-to-3/
- Tempo multi-tenancy: https://grafana.com/docs/tempo/latest/operations/manage-advanced-systems/multitenancy/
- Mimir TLS: https://grafana.com/docs/mimir/latest/manage/secure/securing-communications-with-tls/
- Alloy Promtail migration: https://grafana.com/docs/alloy/latest/set-up/migrate/from-promtail/
- Alloy
otelcol.processor.tail_sampling: https://grafana.com/docs/alloy/latest/reference/components/otelcol/otelcol.processor.tail_sampling/ grafana/otel-lgtmimage: https://github.com/grafana/docker-otel-lgtm