Skip to content

Explanation

Why the Observability 2.0 paradigm exists and how it works: where the term came from, why arbitrarily-wide structured events change the cost and debugging model, how wide events relate to OpenTelemetry signals, why tail sampling is part of the design, what a backend must provide, and where the idea is contested.

Where the other material lives

Step-by-step recipes (middleware code, Collector tail-sampling config, SQL queries, migration steps) are in How-to Guides. Look-up tables (glossary, field groups, sampling keep-rates, backend feature matrix, timeline) are in Reference.

Origins and Timeline

"Observability 2.0" is a label, not a product or a specification. It names an architectural position that Honeycomb and its founders argued for years before the label existed.

  • 2016 — canonical log lines. Brandur Leach described Stripe's canonical log line: in addition to normal logs, each request emits one long line at the end with its key characteristics (brandur.org, 2016-11-26). Stripe's engineering blog restated the pattern in 2019 (Stripe, 2019-07-30).
  • 2016 onward — Honeycomb. Honeycomb (co-founded by Charity Majors and Christine Yen) built its product on a custom distributed column store, Retriever, because it judged existing metrics and log stores unable to answer high-cardinality questions quickly (Why We Built Our Own Distributed Column Store).
  • 2022 — "arbitrarily wide structured events". Majors' essay recommends "arbitrarily wide, structured raw events that are unique and ordered and trace-aware and without any aggregation at write time" (Live Your Best Life With Structured Events).
  • 2023-12 — the versioning idea. Majors floated semantic versioning for observability in a December 2023 post on X.
  • 2024-08-07 — the defining essay. Is It Time To Version Observability? (Signs Point To Yes) sets out 1.0 as "many pillars, many tools" and 2.0 as wide, structured canonical log events as the single source of truth. A version appeared on Honeycomb's blog in late 2024 (Honeycomb).
  • 2025 — the database vendors adopt the language. GreptimeDB positioned itself as "the database for" Observability 2.0 (Greptime, 2025-04-25). ClickHouse bought HyperDX (announced 2025-03-13) and launched ClickStack in May 2025, describing its storage model as wide events (ClickHouse).
  • 2025-10-30 — "the pillar is a lie". Majors argued that the pillar vocabulary itself is vendor marketing that keeps teams in an old mental model (charity.wtf).
  • 2026 — consolidation. Observability Engineering, 2nd Edition (Majors, Fong-Jones, Miranda, with Austin Parker) frames the shift as moving "from collecting separate, disparate signals to unified data workflows" (Honeycomb announcement). In March 2026 Honeycomb itself made time-series Honeycomb Metrics generally available (Honeycomb), which shows that 2.0 in practice means events first, not events only.

See Reference for the same events as a table.

Observability 1.0 vs 2.0 Architecture

This diagram contrasts the two data paths: 1.0 fans one application out to three signal-specific stores, 2.0 funnels one wide event per unit of work into one columnar store that serves every view.

graph TB
    subgraph V1["Observability 1.0 — Three Pillars"]
        direction TB
        APP1["Application"] --> PROM_EXP["Prometheus client / exporter"]
        APP1 --> LOG_AGENT["Fluent Bit / Promtail / Alloy"]
        APP1 --> OTEL_SDK1["OTel SDK (traces)"]

        PROM_EXP --> MIMIR["Prometheus / Mimir"]
        LOG_AGENT --> LOKI["Loki / Elasticsearch"]
        OTEL_SDK1 --> TEMPO["Tempo / Jaeger"]

        MIMIR --> GRAFANA1["Grafana dashboards"]
        LOKI --> GRAFANA1
        TEMPO --> GRAFANA1
    end

    subgraph V2["Observability 2.0 — Unified Wide Events"]
        direction TB
        APP2["Application"] --> WIDE["Wide-event middleware<br/>(enriched OTel spans)"]
        WIDE --> OTEL_COL["OTel Collector<br/>(tail_sampling)"]
        OTEL_COL --> UNIFIED_DB["Columnar event store<br/>(Honeycomb Retriever / ClickHouse / GreptimeDB)"]

        UNIFIED_DB --> DASH["Dashboards and SLOs<br/>(derived at query time)"]
        UNIFIED_DB --> EXPLORE["Exploratory queries<br/>(SQL, BubbleUp-style)"]
        UNIFIED_DB --> ALERT["Alerts and triggers"]
        UNIFIED_DB --> TRACE_VIEW["Trace waterfall"]
    end

    style APP1 fill:#e74c3c,color:#fff
    style APP2 fill:#27ae60,color:#fff
    style UNIFIED_DB fill:#2980b9,color:#fff

Key difference: in 1.0 the application emits three separately shaped signals to three stores, and correlation happens in a human's head or through shared IDs in a UI. In 2.0 the application emits one wide event per request per service hop. Metrics, trace waterfalls and log-style views are derived from that raw record at read time.

Why Wide Events Change the Cost Model

The central technical argument is about where aggregation happens and what that does to cost and to the questions you can ask.

Metrics-first (1.0) Event-first (2.0)
Aggregation point Write time (counters, histograms, recording rules) Read time (GROUP BY over raw events)
Cost driver Number of distinct label combinations (time series) Number and size of events stored
Effect of adding user_id Cardinality explosion: one series per user per metric One more column. Cost barely changes
Questions answerable Only those anticipated when the metric was defined Any combination of fields present on the event
Loss Detail is discarded at write time and cannot be recovered Detail is kept. Volume is controlled by sampling instead

Two properties do most of the work:

  • High cardinality — a field with many unique values (user_id, trace_id, build_sha). These are precisely the fields that locate a specific customer, deploy or shard during an incident. Time-series databases key storage on label combinations, so every new high-cardinality label multiplies series count.
  • High dimensionality — many fields per event. Each extra field is another axis you can slice by later without re-instrumenting. Columnar stores only read the columns a query touches, so an event with 200 fields costs little more to query than one with 20.

Majors' summary of the consequence: 1.0 tools "can only give you aggregates and random exemplars", while 2.0 tools can tell you precisely what happened when you flipped a flag or deployed a canary (charity.wtf, 2024).

Pre-aggregation does not disappear

Removing metrics as first-class instrumentation does not remove pre-aggregation. It moves it from the application (and the metric SDK) into the database, as materialized views or continuous aggregation computed from raw events.

Wide Event Anatomy

A wide event captures the complete context of one unit of work (usually one request in one service) in one structured record. This example from a checkout service shows about 30 fields across 7 context groups:

{
  "timestamp": "2025-01-15T10:23:45.612Z",
  "request_id": "req_8bf7ec2d",
  "trace_id": "abc123",

  "service": "checkout-service",
  "version": "2.4.1",
  "deployment_id": "deploy_789",
  "region": "us-east-1",

  "method": "POST",
  "path": "/api/checkout",
  "status_code": 500,
  "duration_ms": 1247,

  "user": {
    "id": "user_456",
    "subscription": "premium",
    "account_age_days": 847,
    "lifetime_value_cents": 284700
  },

  "cart": {
    "id": "cart_xyz",
    "item_count": 3,
    "total_cents": 15999,
    "coupon_applied": "SAVE20"
  },

  "payment": {
    "method": "card",
    "provider": "stripe",
    "latency_ms": 1089,
    "attempt": 3
  },

  "error": {
    "type": "PaymentError",
    "code": "card_declined",
    "message": "Card declined by issuer",
    "retriable": false,
    "stripe_decline_code": "insufficient_funds"
  },

  "feature_flags": {
    "new_checkout_flow": true,
    "express_payment": false
  }
}

When a user complains, filtering on user.id = "user_456" immediately shows:

  • Premium customer with a 2+ year account (high priority)
  • Payment failed on the 3rd attempt with insufficient_funds
  • The new checkout flow was enabled (a correlation worth checking across all failures)
  • No grep, no guessing, no second search

The field groups (identity, infrastructure, request, user/business, operation, error, experiments) are listed in Reference. Example queries against events like this are in How-to Guides.

Structured logging is necessary but not sufficient

Structured logging means JSON instead of free text, which is table stakes. Wide events are a discipline: one complete event per unit of work, with all context attached. Structured logs can still be useless: 5 fields, no user context, scattered across 20 lines per request.

Wide Event Lifecycle

The implementation idea is simple: create the event when the request starts, let every layer add fields as it learns things, and emit exactly once at the end. The sequence below shows that lifecycle for the checkout example; the runnable code is in How-to Guides.

sequenceDiagram
    participant Client
    participant Middleware as Wide-event middleware
    participant Handler as Checkout handler
    participant Deps as Cart DB / payment provider
    participant Emitter as Logger or OTel span
    participant Backend as Event store

    Client->>Middleware: POST /api/checkout
    Middleware->>Middleware: Create event (request_id, method, path, service, version)
    Middleware->>Handler: next()
    Handler->>Handler: Add user.id, user.subscription, lifetime value
    Handler->>Deps: getCart(user.id)
    Deps-->>Handler: cart
    Handler->>Handler: Add cart.id, item_count, total_cents
    Handler->>Deps: processPayment(cart, user)
    Deps-->>Handler: payment result or decline
    Handler->>Handler: Add payment.provider, latency_ms, attempt, error.*
    Handler-->>Middleware: response
    Middleware->>Middleware: Finalize duration_ms, status_code, outcome
    Middleware->>Emitter: Emit one wide event
    Emitter->>Backend: One record, about 30 to 50 fields
    Middleware-->>Client: HTTP response

Wide Events and OpenTelemetry

OpenTelemetry (OTel) is how most teams will transport wide events, but OTel's own model and vocabulary differ from the Observability 2.0 framing. Keeping the terms apart avoids confusion (see also OpenTelemetry and OTel Collector).

OTel's Stance: Separate Signals, Shared Context

OTel defines distinct signals (traces, metrics, logs, and the newer profiles signal) and makes them correlatable through shared context (trace and span IDs) and shared resource attributes. An OTel maintainer post states the design directly: "all of your telemetry signals can be correlated through context" and "OpenTelemetry is built on the concept that all signals are interpreted together, rather than separately" (OpenTelemetry Logging and You, 2025-04-18). OTel does not use the "Observability 2.0" label and does not prescribe a single storage model. That choice is left to backends.

Terminology Clash: "Wide Event" vs OTel "Event"

Term Meaning Has duration?
Wide event (Observability 2.0) One arbitrarily-wide record per unit of work, typically per request per service. In OTel terms it is usually a span enriched with many attributes Yes (duration_ms)
OTel Event A LogRecord with a mandatory event_name and a schema defined by semantic conventions, emitted via the Logs API. "Not all logs are events, but all events are logs" No
Span event A timestamped annotation on a span (Span.AddEvent, Span.RecordException) No

The same OTel post notes that a span is essentially "an event with a particularly detailed schema", distinguished by having a duration and a parent-child hierarchy. That is why Honeycomb-style practice treats enriched spans as the wide events: one span per service hop carries the business context, and the trace links the hops.

Recent OTel Changes That Help Wide Events

  • Complex attribute types on all signals. Following OTEP 4485 and OTLP 1.9.0, OTel APIs and SDKs are adding maps, heterogeneous arrays, byte arrays and empty values as attribute values on every signal, not only logs. The OTel authors (including Honeycomb's Austin Parker) caution that many backends cannot index or aggregate complex attributes, and advise primitive values where possible (OTel blog, 2025-11-05). In practice, flattened dotted keys (app.user.id) remain the portable way to build wide spans.
  • Span Event API deprecation. OTel is deprecating Span.AddEvent and Span.RecordException in favor of log-based events correlated with the active span; existing span-event data and views keep working through a compatibility layer (OTel blog, 2026-03-17, OTEP 4430). The Events API/SDK had already been marked deprecated in favor of the Logs API with event_name. For wide-event practice this reinforces the rule: put request context on the span attributes, not in span events.

OpenTelemetry alone does not make telemetry wide

OTel standardizes how telemetry is produced and shipped. It does not decide what context to capture. Auto-instrumentation yields spans with a name, duration, status and HTTP/DB semantic-convention attributes, but no user, cart, plan or feature-flag context. Without deliberate enrichment you get narrow telemetry in a standard format.

When Teams Bypass OTel

The OTel pipeline is not always the cheapest path for very high volumes of wide events. ClickHouse reported that its internal LogHouse platform grew beyond 100 PB of uncompressed data, and that it replaced OTel for ClickHouse's own system logs with a purpose-built exporter (SysEx) that preserves native ClickHouse types and avoids intermediate conversions (ClickHouse, June 2025). This is a vendor case study for one specialized source, not a general recommendation to drop OTel.

Sampling in an Event-First World

Storing every raw event is the ideal. At high volume it is not affordable, so 2.0 systems control cost with sampling instead of pre-aggregation.

Head vs Tail Sampling

  • Head sampling decides at the start of a request (usually by trace ID). It is cheap and stateless, but blind: it drops a random 99% whether the request later fails or not.
  • Tail sampling decides after the request (or whole trace) completes, using its outcome: status, latency, customer tier, feature flag. It keeps the rare, interesting events and thins the common, healthy ones.

Naive random sampling is dangerous

Sampling 1% of all traffic at random can drop the one request that explains an outage. Use outcome-based rules and keep 100% of errors and slow requests.

Sample Rates Must Travel With the Event

Tail sampling skews the stored data on purpose (all errors, few successes). A backend can only compute correct counts and rates if each kept event records the rate it was sampled at, so a query can weight it (an event kept at 1-in-20 counts as 20). Honeycomb's Refinery proxy and its dynamic sampler work this way (Refinery README). If your pipeline does not propagate sample rates, derived error rates will be wrong.

Where Tail Sampling Runs

The diagram shows the standard two-tier OpenTelemetry Collector layout: agents route every span of a trace to the same gateway replica, because the tail_sampling processor can only decide once it holds the whole trace in memory.

sequenceDiagram
    participant SDK as App with OTel SDK
    participant Agent as Collector agent<br/>(loadbalancing exporter)
    participant GW as Collector gateway<br/>(tail_sampling processor)
    participant Store as Event store

    SDK->>Agent: OTLP spans (enriched with app.* attributes)
    Agent->>GW: Route by traceID so one replica sees the whole trace
    GW->>GW: Buffer spans for decision_wait (default 30s)
    GW->>GW: Evaluate policies (status_code, latency, string_attribute, probabilistic)
    alt Policy matched
        GW->>Store: Export all spans of the trace
    else No policy matched
        GW->>GW: Drop trace, remember decision in cache
    end

The processor is beta for traces and ships in the contrib and k8s Collector distributions. Its README warns that it is stateful: all spans of a trace must reach the same instance, which is why the agent tier uses the loadbalancing exporter (tail sampling processor README). Memory sizing and placement are covered in OTel Collector tail sampling placement. The keep-rate table is in Reference and the config is in How-to Guides.

Database Requirements for Observability 2.0

Wide events are large. A single uncompressed wide event can exceed 2 KB (Greptime, 2025-04-25); at 10,000 requests/second that is about 20 MB/s of raw event data before compression. The store must ingest that continuously and serve both real-time dashboards and ad-hoc exploration.

The diagram shows the logical pipeline such a store implements, independent of vendor.

graph LR
    subgraph Ingest
        OTLP["OTLP receiver"]
        TRANSFORM["Ingest-time transform<br/>(parse, enrich, flatten)"]
    end

    subgraph Store
        COLUMNAR["Columnar parts<br/>(per-column encoding and compression)"]
        OBJECT["Object storage<br/>(S3 / GCS / Azure Blob)"]
        MATVIEW["Materialized views /<br/>continuous aggregation"]
    end

    subgraph Query
        ROUTINE["Routine queries<br/>(dashboards, alerts, SLOs)"]
        EXPLORE["Exploratory queries<br/>(ad-hoc SQL, group-by-anything)"]
        PROMQL["PromQL<br/>(backward compatibility)"]
    end

    OTLP --> TRANSFORM --> COLUMNAR
    COLUMNAR --> OBJECT
    COLUMNAR --> MATVIEW
    COLUMNAR --> EXPLORE
    MATVIEW --> ROUTINE
    MATVIEW --> PROMQL

    style COLUMNAR fill:#2980b9,color:#fff
    style MATVIEW fill:#8e44ad,color:#fff
Requirement Why How backends implement it
Columnar storage Wide events have 50+ fields; queries touch few. Column pruning and vectorized execution avoid reading the rest ClickHouse MergeTree parts, GreptimeDB Parquet-based SSTs, Honeycomb Retriever segments
Disaggregated compute and storage Retention grows faster than query load; storage must scale independently Object storage as the primary tier, with local disk or memory caches for recent data
Dynamic schema New fields appear whenever instrumentation changes; an ALTER TABLE per attribute does not scale Auto-created columns (GreptimeDB), the JSON type with per-path subcolumns (ClickHouse, production-ready since 25.3), columns created on first appearance of a field (Honeycomb)
High-cardinality filtering user_id, trace_id, request_id have millions of values Inverted, skipping and bloom-filter indexes; sort keys on common filters
Fresh data Dashboards and alerts need data within seconds WAL plus memtable designs; streaming inserts
Materialized views Error rates and p99s must be cheap for dashboards Incremental aggregation that updates without reprocessing raw events
PromQL compatibility Existing Grafana dashboards and alert rules should keep working A PromQL engine over the columnar store (GreptimeDB), or metrics tables queried from Grafana
Workload isolation Heavy exploration must not starve alerting Read replicas or separate compute pools (GreptimeDB Enterprise, ClickHouse Cloud compute-compute separation)

Routine vs Exploratory Queries

Query type Purpose Latency target Example
Routine Dashboards, alerts, SLO tracking Sub-second Error rate by service over the last 5 minutes
Exploratory Ad-hoc debugging of unknown unknowns Seconds to minutes "All requests from user X where flag Y was on and latency > 2 s"

Routine queries are predictable and can be served from materialized views. Exploratory queries are unpredictable and must scan raw events, which is why they are the workload that separates 2.0 backends from 1.0 ones.

Backend Architectures

Three backends illustrate the design space. Their feature facts are tabulated in Reference.

Honeycomb Retriever

Honeycomb's SaaS stores all customer events in Retriever, a custom distributed column store. Events arrive through Kafka and are written column-by-column to disk; queries fan out and scan only the columns they need, which is what makes group-by on any field fast. Honeycomb's pitch is the query experience on top: BubbleUp compares a selected slice of events against the baseline to show which dimensions differ (Why Observability Requires a Distributed Column Store). Tail sampling can happen before ingestion in Refinery (Apache-2.0).

ClickHouse and ClickStack

ClickStack (launched May 2025) bundles ClickHouse, the HyperDX UI (MIT-licensed, acquired by ClickHouse in March 2025) and an OpenTelemetry Collector configured with an opinionated ClickHouse schema. The project states that "all observability data should be ingested as wide, rich events", but those events are "stored in ClickHouse tables by data type - logs, traces, metrics, and sessions" and correlated at query time (ClickStack README). So ClickStack is event-first storage with per-signal tables, not a single table of wide events. ClickHouse also underpins other OTel-native platforms such as SigNoz.

GreptimeDB Reference Architecture

GreptimeDB is an Apache-2.0 (core), Rust-based observability database that runs metrics, logs and traces "on one columnar engine over object storage" with one table model of tags, timestamp and fields (GreptimeDB README). The diagram shows its distributed mode using the project's own component names.

graph TB
    subgraph Ingestion
        OTLP_IN["OTLP"]
        PROM_RW["Prometheus Remote Write"]
        OTHER_IN["Loki Push / Elasticsearch Bulk /<br/>InfluxDB line protocol"]
    end

    subgraph Cluster["GreptimeDB distributed mode"]
        FRONTEND["Frontend<br/>(protocol entry, distributed query engine, stateless)"]
        METASRV["Metasrv<br/>(metadata, routing, repartitioning)"]
        DATANODE["Datanode<br/>(regions: WAL, memtable, SST, indexes, compaction)"]
        FLOWNODE["Flownode (optional)<br/>(continuous aggregation / materialized views)"]
        KV["etcd or RDS<br/>(metadata KV)"]
    end

    subgraph Storage
        CACHE["Memory and local-disk cache<br/>(recent, hot data)"]
        S3["Object storage<br/>(S3 / GCS / Azure Blob)"]
    end

    subgraph Consumers
        GRAFANA["Grafana<br/>(PromQL, SQL)"]
        SQL_CLIENT["SQL clients<br/>(MySQL / PostgreSQL wire)"]
        JAEGER["Jaeger-compatible trace UI"]
    end

    OTLP_IN --> FRONTEND
    PROM_RW --> FRONTEND
    OTHER_IN --> FRONTEND
    FRONTEND --> DATANODE
    FRONTEND --> METASRV
    METASRV --> KV
    DATANODE --> CACHE
    DATANODE --> S3
    DATANODE --> FLOWNODE
    FLOWNODE --> FRONTEND

    FRONTEND --> GRAFANA
    FRONTEND --> SQL_CLIENT
    FRONTEND --> JAEGER

    style FRONTEND fill:#2980b9,color:#fff
    style DATANODE fill:#27ae60,color:#fff
    style FLOWNODE fill:#8e44ad,color:#fff
    style S3 fill:#e67e22,color:#fff

Key GreptimeDB features for Observability 2.0, per the project README and docs:

  • OTLP ingestion — spans land in opentelemetry_traces and log records in opentelemetry_logs, both carrying trace_id, so correlation is a SQL join
  • Pipelines — ingest-time parsing and transformation
  • Flow engine — continuous aggregation (streaming and materialized views) to derive metrics from raw events
  • Automatic columns and JSON2 — new attributes become columns; v1.2.0 (2026-09-08) added a structural JSON type stored as structs with dot-style SQL access
  • PromQL, SQL and Jaeger-compatible queries, plus MySQL and PostgreSQL wire protocols
  • Edition boundary — cluster mode, object storage and Flow are in the Apache-2.0 build; read replicas, workload isolation and automated repartitioning are Enterprise-only

Correction (2026-09)

Earlier versions of this page drew GreptimeDB as "Ingest Nodes", "Query Nodes" and a "Rule Engine". Those are not GreptimeDB component names; the real distributed components are Frontend, Datanode, Metasrv and Flownode. Read replicas are an Enterprise feature, not part of the open-source core.

Common Misconceptions

Misconception Reality
"Structured logging is the same as wide events" Structured logging is JSON instead of strings. Wide events are a discipline: one event per unit of work with all context attached
"We already use OTel, so we're good" OTel is a delivery mechanism. Auto-instrumented spans carry name, duration, status and protocol attributes. Business context must be added deliberately
"This is just tracing with extra steps" Tracing shows request flow across services. Wide events carry the context within each service. Ideally your wide events are your spans, enriched
"Logs are for debugging, metrics are for dashboards" That split is an artifact of storage engines. Wide events serve both: query them to debug, aggregate them for dashboards
"High-cardinality data is expensive and slow" It is expensive in time-series databases and in logging systems built for full-text search. Columnar stores (ClickHouse, GreptimeDB, Honeycomb Retriever) are built to filter and group on it
"2.0 means no metrics" Even Honeycomb added time-series metrics (GA March 2026). Infrastructure counters such as CPU or queue depth have no request to attach to and remain metrics

Criticisms and Limits

The paradigm is influential but not settled. The strongest objections, including from its proponents:

  • The version number is marketing. Hazel Weakly, who proposed an "Observability 3.0" framing, has called the 1.0/2.0/3.0 versioning "entirely marketing", and Majors has said she dislikes the framing while keeping it for its explanatory power (Another observability 3.0 appears on the horizon, 2025-03-24; Hazel Weakly).
  • Vendor alignment. The chief advocates sell event stores (Honeycomb, Greptime, ClickHouse). Their cost comparisons are vendor-authored; independent, like-for-like cost studies at scale remain scarce (see the index open questions).
  • Not everything is a request. Host metrics, queue depths, garbage-collection counters and batch jobs do not map onto one-event-per-request; they are cheaper as metrics.
  • Sampling reintroduces loss. Event-first systems still discard data at high volume. Getting weighting and rules wrong silently skews rates.
  • Governance risk. Business context (user IDs, plan tiers, sometimes e-mail or payment data) turns telemetry into a store of personal data. Wide events need attribute allow-lists, redaction and retention rules that metrics never needed.
  • Unified storage is often per-signal in practice. ClickStack and GreptimeDB's OTLP ingestion both keep traces and logs in separate tables and correlate them with joins, so "single source of truth" usually means one engine and shared IDs, not literally one table.

Sources