Observability¶
Summary
Telemetry standards, collection pipelines, storage backends, visualization, APM platforms and eBPF tooling for keeping cloud-native systems observable. The domain spans the instrumentation standard (OpenTelemetry, now CNCF Graduated) and its pipeline engine (OTel Collector), two composable backend stacks (LGTM with Grafana, and the Victoria Stack), five all-in-one platforms (SigNoz, OpenObserve, Coroot, Apache SkyWalking, Monoscope), the Observability 2.0 wide-events paradigm, and kernel-level tooling (bpftrace, eBPF Developer Tutorial). Versions and licenses below match the topic pages as of 2026-09-25.
Domain Map¶
The map shows how the topics relate: OpenTelemetry defines the data, collectors move it, backends store it, and eBPF tooling works at the kernel layer underneath.
flowchart LR
subgraph STD["Standard and pipeline"]
OTEL["OpenTelemetry<br/>(spec 1.61, OTLP)"]
COL["OTel Collector 0.161<br/>(Alloy is a distribution)"]
end
subgraph COMP["Composable stacks"]
LGTM["LGTM: Mimir, Loki,<br/>Tempo, Pyroscope"]
VIC["Victoria Stack:<br/>VictoriaMetrics, Logs, Traces"]
GRAF["Grafana 13"]
end
subgraph ALL["All-in-one platforms"]
SIG["SigNoz<br/>(ClickHouse)"]
O2["OpenObserve<br/>(Parquet on S3)"]
SW["SkyWalking<br/>(OAP + BanyanDB)"]
MONO["Monoscope<br/>(TimescaleDB, TimeFusion)"]
COR["Coroot<br/>(eBPF + ClickHouse)"]
end
subgraph KERN["Kernel layer"]
BPFT["bpftrace"]
TUT["eBPF Developer Tutorial"]
end
O20["Observability 2.0<br/>(wide events)"]
OTEL -->|"OTLP"| COL
COL --> LGTM
COL --> VIC
COL --> SIG
COL --> O2
COL --> SW
COL --> MONO
GRAF -.->|"queries"| LGTM
GRAF -.->|"queries"| VIC
TUT -.->|"teaches the hooks used by"| COR
BPFT -.->|"same kernel probes"| COR
O20 -.->|"columnar event stores"| SIG
Topics¶
| Topic | Category | Highlight | Latest version (2026-09-25) | License |
|---|---|---|---|---|
| OpenTelemetry | Standard | Vendor-neutral APIs, SDKs, OTLP and semantic conventions; CNCF Graduated (2026-05). Topic covers the official blueprints (Kubernetes and non-Kubernetes) and the Adobe, Mastodon and Skyscanner reference implementations | Spec v1.61.0 (2026-09-14); Operator v0.159.0 | Apache-2.0 |
| OTel Collector | Pipeline | Receivers, processors, exporters, connectors; six official distributions, agent/gateway topologies, tail sampling, two-week releases with no LTS | v1.67.0 / v0.161.0 (2026-09-14) | Apache-2.0 |
| Grafana | Visualization | Dashboards and alerting over 100+ data sources; Git Sync and dynamic dashboards GA in 13.0; Grafana Assistant | 13.2.2 (2026-09-15) | AGPL-3.0 |
| LGTM Stack | Composable stack | Loki, Grafana, Tempo, Mimir plus Pyroscope and Alloy; object storage, Kafka-decoupled write paths (Mimir 3, Tempo 3) | Mimir 3.2.1, Loki 3.7.8, Tempo 3.0.3, Pyroscope 2.3.1, Alloy 1.20.0 | AGPL-3.0 (Alloy Apache-2.0) |
| Victoria Stack | Composable stack | VictoriaMetrics, VictoriaLogs, VictoriaTraces: single binaries on local disks, no external dependencies | VictoriaMetrics v1.152.0 (2026-09-14), VictoriaLogs v1.52.0, VictoriaTraces v0.11.1 (pre-1.0) | Apache-2.0 (Enterprise commercial) |
| SigNoz | All-in-one | OpenTelemetry-native logs, metrics, traces, APM and LLM observability on ClickHouse; installs via Foundry or Helm | v0.143.0 (2026-09-23) | MIT core, ee/ commercial, collector AGPL-3.0 |
| OpenObserve | All-in-one | Rust platform storing logs, metrics, traces, RUM, profiles and LLM traces as Parquet on object storage; SQL and PromQL; 1.0 GA in 2026-09 | v1.0.4 (2026-09-24) | AGPL-3.0 (Enterprise commercial) |
| Coroot | All-in-one (eBPF) | Zero-instrumentation eBPF metrics, logs, traces and profiles; service map, inspections, SLOs; AI RCA and MCP endpoint | v1.26.8 (2026-09-24) | Apache-2.0 (Enterprise commercial) |
| Apache SkyWalking | All-in-one (APM) | ASF APM with language agents, Rover eBPF, mesh telemetry; OAP server on BanyanDB; Horizon UI since 11.0 | OAP 11.0.0 (2026-08-28), BanyanDB 0.11.1 | Apache-2.0 |
| Monoscope | All-in-one (AI) | OTel ingest, natural-language-to-KQL search, scheduled AI agents, MCP. TimescaleDB is still the primary store; the TimeFusion (S3) migration is in progress | v0.6.27 (2026-09-07), pre-1.0 | AGPL-3.0 (TimeFusion MIT) |
| Observability 2.0 | Paradigm | Wide structured events as the single source of truth instead of three pillar silos; columnar stores and tail sampling | N/A (concept) | N/A |
| bpftrace | eBPF tooling | awk-like tracing language compiled through LLVM and loaded via libbpf; one-liners, per-CPU aggregation; kernel 6.1+ with BTF | 0.27.0 (2026-09-10) | Apache-2.0 |
| eBPF Developer Tutorial | eBPF learning | CO-RE eBPF course: 65 runnable tutorials (23 executed in CI) from kprobe basics to sched_ext, GPU tracing and Linux 7.0 features | No releases; last commit 2026-09-05 | MIT |
Comparisons¶
| Comparison | Scope |
|---|---|
| Observability Stacks Comparison | Coroot, SigNoz, SkyWalking, OpenObserve, LGTM, Victoria and Monoscope: signals, architecture, ingestion, reliability, cost, security, with a decision flowchart |
| LGTM vs Victoria Stack | Head-to-head of the two composable stacks, signal by signal, with migration paths and a decision flowchart |
| eBPF Tutorial vs libbpf-bootstrap vs bpftrace | Layer-mapped comparison of a learning curriculum, a development scaffold and an interactive tracing language |
When to Use Which¶
| Need | Reach for |
|---|---|
| Vendor-neutral instrumentation for any backend | OpenTelemetry SDKs + OTel Collector |
| Dashboards over many data sources | Grafana |
| Full metrics, logs, traces and profiles with deep correlation, and a platform team | LGTM Stack |
| Lowest resource use, local disks, Apache-2.0 | Victoria Stack |
| One OTel-native product and UI, Datadog replacement | SigNoz |
| Elasticsearch replacement for logs on object storage, with SQL | OpenObserve (AGPL) or VictoriaLogs (Apache-2.0) |
| Observability without code changes, SLO alerts, root-cause hints | Coroot |
| Java-heavy estates, Istio/Envoy meshes, ASF governance | Apache SkyWalking |
| API-heavy product teams wanting AI-assisted search and reports (pre-1.0) | Monoscope |
| LLM and agent application monitoring | SigNoz, OpenObserve or SkyWalking (Grafana Cloud as a managed option) |
| Ad-hoc kernel or process investigation now | bpftrace |
| Learning to write and ship eBPF tools | eBPF Developer Tutorial, then libbpf-bootstrap or a starter template |
The stacks comparison has the full decision flowchart.
Landscape¶
The space is converging on OpenTelemetry as the universal instrumentation standard. OpenTelemetry became a CNCF Graduated project in May 2026 and is one of the most active CNCF projects. It provides one set of SDKs, the OTLP protocol and a Collector for metrics, logs and traces, with profiles in public alpha since 2026-03. It replaces the fragmented mix of Prometheus client libraries, Jaeger SDKs and Fluentd/Fluent Bit log agents, although those remain common as receivers.
eBPF-based auto-instrumentation is a zero-code alternative. Coroot, Grafana Beyla / OpenTelemetry eBPF Instrumentation (OBI) and Odigos extract HTTP, gRPC and database telemetry from kernel-level events without modifying application code, and SkyWalking Rover adds eBPF on/off-CPU and network profiling. The Observability 2.0 paradigm challenges the "three pillars" model and argues for wide structured events as the single source of truth. Profiling has become a fourth signal: Grafana Pyroscope, Coroot, SkyWalking, OpenObserve and Parca provide continuous profiling, and OTLP now has a profiles signal.
Storage economics
Object storage (S3, GCS, Azure Blob) is the default persistence layer for many modern backends: Mimir, Loki, Tempo and Pyroscope keep long-term data there, OpenObserve writes Parquet files to it, Coroot can tier ClickHouse data to S3, and Monoscope's TimeFusion writes Delta Lake tables to a bucket. This makes long retention mainly a cost setting. The Victoria Stack is the deliberate counterexample: it keeps data on local disks and uses object storage only for backups, trading unbounded retention for fewer dependencies and lower query latency.
The market splits between all-in-one platforms (SigNoz, OpenObserve, Coroot, SkyWalking, Monoscope) that store all signals in one product, and composable stacks (LGTM, Victoria) that optimize each signal separately. In 2025-2026 several backends grew heavier write paths (Kafka in Mimir 3 and Tempo 3) while others shed dependencies.
AI features became standard quickly: natural-language querying (Monoscope, Grafana Assistant, OpenObserve's AI Assistant), automated root-cause analysis (Coroot Enterprise, Grafana Assistant Investigations), MCP endpoints for AI agents (Coroot, SigNoz, SkyWalking Horizon, Monoscope, Grafana Cloud) and LLM/agent observability (SigNoz, OpenObserve, SkyWalking, Grafana Cloud).
Key Concepts¶
Three Signals (Metrics, Logs, Traces)¶
Signal types
- Metrics: Numeric time series (counters, gauges, histograms) sampled at regular intervals. Low cardinality, high compression, ideal for alerting and dashboards. Prometheus exposition format and OpenTelemetry metrics are the two dominant formats.
- Logs: Timestamped records of discrete events. High volume, semi-structured, essential for debugging. Loki indexes only stream labels and compresses log lines in chunks; VictoriaLogs stores every field as a column with bloom filters.
- Traces: Distributed call graphs made of spans, each a unit of work with timing, status and parent-child links. Critical for latency in microservice architectures.
Observability 2.0 alternative
The Observability 2.0 paradigm unifies these signals into wide structured events: one rich record per request, from which metrics, logs and traces are derived at query time.
OpenTelemetry¶
A CNCF Graduated project (since 2026-05) providing vendor-neutral APIs, SDKs, OTLP and the OTel Collector for generating, collecting, processing and exporting telemetry. The Collector is a pipeline with three stages:
- Receivers: ingest OTLP, Prometheus scrape, Jaeger, Zipkin, Fluent Forward, syslog and dozens of other protocols (most live in the contrib distribution).
- Processors: transform data in flight: batching, head or tail sampling, attribute changes, filtering, Kubernetes metadata.
- Exporters: send data to any backend: OTLP to Mimir, Tempo, Loki or SigNoz, Prometheus remote write, Kafka, ClickHouse or SaaS platforms. Connectors (for example span-to-metrics) join pipelines.
The tail-sampling processor examines complete traces before deciding what to keep, which cuts storage cost and preserves error and slow traces.
Component renames
Since v0.144.0 the otlp and otlphttp exporters are named otlp_grpc and otlp_http, with deprecated aliases, as part of a snake_case rename campaign. The Collector ships every two weeks with no LTS. See OTel Collector.
For the official best-practice corpus (blueprints versus reference implementations, and how Adobe, Mastodon and Skyscanner run collectors in production) see the OpenTelemetry topic.
Cardinality¶
The number of unique time series created by combinations of metric name and label values. High cardinality (labels such as user IDs, request paths or container IDs) drives storage growth and slows queries. Prometheus and Mimir enforce per-tenant series limits. VictoriaMetrics is reported to use several times less RAM per series and offers optional limits (-storage.maxHourlySeries, -storage.maxDailySeries). Controls include recording rules (pre-aggregation), relabeling (dropping labels), Loki structured metadata instead of labels, and OTel Collector attribute processors. Column stores (ClickHouse in SigNoz and Coroot, Parquet in OpenObserve) handle high-cardinality attributes in logs and traces better than label indexes.
Exemplars¶
Exemplars link a metric sample to a representative trace span, bridging aggregates and individual requests. When a histogram bucket records a slow request, the exemplar carries its trace ID, so a dashboard click opens the exact trace. Grafana overlays exemplars on metric panels, and Mimir stores them. VictoriaMetrics does not store exemplars (a long-open feature request); Grafana correlations are the workaround.
SLI/SLO (Service Level Indicators and Objectives)¶
SLIs are quantitative measures of service behavior (for example, the share of requests completing under 300 ms). SLOs set targets for SLIs over a rolling window (for example, 99.9% of requests succeed over 30 days). Error budgets, the allowed unreliability, drive decisions: when the budget is spent, teams slow deployments and focus on reliability. Sloth, Pyrra and OpenSLO generate Prometheus recording and burn-rate alerting rules from SLO definitions. Coroot and OpenObserve have built-in SLO alerting.
Related¶
- Kubernetes: where the Operator, collectors and most backends run
- Istio and Cilium: mesh telemetry for SkyWalking and OTel, eBPF networking alongside Coroot
- Apache Kafka and NATS: Kafka in the Mimir 3 and Tempo 3 write paths; NATS as OpenObserve's cluster coordinator
- MinIO: S3-compatible storage for LGTM, OpenObserve and Monoscope TimeFusion
- PostgreSQL: metadata store for Grafana, OpenObserve, SigNoz (optional) and Monoscope
- Argo CD: GitOps delivery of Collector resources and dashboards
- AI Platform Engineering: inference telemetry built on OTel conventions
Sources¶
- OpenTelemetry and CNCF graduation announcement
- OpenTelemetry Collector repository
- Grafana documentation (Grafana, Mimir, Loki, Tempo, Pyroscope, Alloy)
- VictoriaMetrics documentation
- SigNoz documentation
- OpenObserve documentation
- Coroot documentation
- Apache SkyWalking
- Monoscope and repository
- Charity Majors: Is it time to version observability? (Observability 2.0)
- bpftrace and eBPF Developer Tutorial
Open Questions¶
- As eBPF auto-instrumentation (OBI, Coroot, Odigos) matures, will manual OpenTelemetry SDK instrumentation shrink to business-specific spans and custom metrics?
- Is aggressive pre-aggregation (recording rules) still the right cardinality trade-off, or do column stores and VictoriaMetrics make raw high-cardinality storage economical?
- With profiling now a fourth signal (OTLP profiles in alpha), what is the realistic production overhead of continuous profiling, and when will OTLP profiles reach beta with broad backend support?
- Is there an independent VictoriaLogs vs Loki vs OpenObserve benchmark at 100 GB/day or more?
- When will Monoscope finish moving telemetry from TimescaleDB to TimeFusion, and will VictoriaTraces reach 1.0?