Explanation¶
Why the Observability 2.0 paradigm exists and how it works: where the term came from, why arbitrarily-wide structured events change the cost and debugging model, how wide events relate to OpenTelemetry signals, why tail sampling is part of the design, what a backend must provide, and where the idea is contested.
Where the other material lives
Step-by-step recipes (middleware code, Collector tail-sampling config, SQL queries, migration steps) are in How-to Guides. Look-up tables (glossary, field groups, sampling keep-rates, backend feature matrix, timeline) are in Reference.
Origins and Timeline¶
"Observability 2.0" is a label, not a product or a specification. It names an architectural position that Honeycomb and its founders argued for years before the label existed.
- 2016 — canonical log lines. Brandur Leach described Stripe's canonical log line: in addition to normal logs, each request emits one long line at the end with its key characteristics (brandur.org, 2016-11-26). Stripe's engineering blog restated the pattern in 2019 (Stripe, 2019-07-30).
- 2016 onward — Honeycomb. Honeycomb (co-founded by Charity Majors and Christine Yen) built its product on a custom distributed column store, Retriever, because it judged existing metrics and log stores unable to answer high-cardinality questions quickly (Why We Built Our Own Distributed Column Store).
- 2022 — "arbitrarily wide structured events". Majors' essay recommends "arbitrarily wide, structured raw events that are unique and ordered and trace-aware and without any aggregation at write time" (Live Your Best Life With Structured Events).
- 2023-12 — the versioning idea. Majors floated semantic versioning for observability in a December 2023 post on X.
- 2024-08-07 — the defining essay. Is It Time To Version Observability? (Signs Point To Yes) sets out 1.0 as "many pillars, many tools" and 2.0 as wide, structured canonical log events as the single source of truth. A version appeared on Honeycomb's blog in late 2024 (Honeycomb).
- 2025 — the database vendors adopt the language. GreptimeDB positioned itself as "the database for" Observability 2.0 (Greptime, 2025-04-25). ClickHouse bought HyperDX (announced 2025-03-13) and launched ClickStack in May 2025, describing its storage model as wide events (ClickHouse).
- 2025-10-30 — "the pillar is a lie". Majors argued that the pillar vocabulary itself is vendor marketing that keeps teams in an old mental model (charity.wtf).
- 2026 — consolidation. Observability Engineering, 2nd Edition (Majors, Fong-Jones, Miranda, with Austin Parker) frames the shift as moving "from collecting separate, disparate signals to unified data workflows" (Honeycomb announcement). In March 2026 Honeycomb itself made time-series Honeycomb Metrics generally available (Honeycomb), which shows that 2.0 in practice means events first, not events only.
See Reference for the same events as a table.
Observability 1.0 vs 2.0 Architecture¶
This diagram contrasts the two data paths: 1.0 fans one application out to three signal-specific stores, 2.0 funnels one wide event per unit of work into one columnar store that serves every view.
graph TB
subgraph V1["Observability 1.0 — Three Pillars"]
direction TB
APP1["Application"] --> PROM_EXP["Prometheus client / exporter"]
APP1 --> LOG_AGENT["Fluent Bit / Promtail / Alloy"]
APP1 --> OTEL_SDK1["OTel SDK (traces)"]
PROM_EXP --> MIMIR["Prometheus / Mimir"]
LOG_AGENT --> LOKI["Loki / Elasticsearch"]
OTEL_SDK1 --> TEMPO["Tempo / Jaeger"]
MIMIR --> GRAFANA1["Grafana dashboards"]
LOKI --> GRAFANA1
TEMPO --> GRAFANA1
end
subgraph V2["Observability 2.0 — Unified Wide Events"]
direction TB
APP2["Application"] --> WIDE["Wide-event middleware<br/>(enriched OTel spans)"]
WIDE --> OTEL_COL["OTel Collector<br/>(tail_sampling)"]
OTEL_COL --> UNIFIED_DB["Columnar event store<br/>(Honeycomb Retriever / ClickHouse / GreptimeDB)"]
UNIFIED_DB --> DASH["Dashboards and SLOs<br/>(derived at query time)"]
UNIFIED_DB --> EXPLORE["Exploratory queries<br/>(SQL, BubbleUp-style)"]
UNIFIED_DB --> ALERT["Alerts and triggers"]
UNIFIED_DB --> TRACE_VIEW["Trace waterfall"]
end
style APP1 fill:#e74c3c,color:#fff
style APP2 fill:#27ae60,color:#fff
style UNIFIED_DB fill:#2980b9,color:#fff
Key difference: in 1.0 the application emits three separately shaped signals to three stores, and correlation happens in a human's head or through shared IDs in a UI. In 2.0 the application emits one wide event per request per service hop. Metrics, trace waterfalls and log-style views are derived from that raw record at read time.
Why Wide Events Change the Cost Model¶
The central technical argument is about where aggregation happens and what that does to cost and to the questions you can ask.
| Metrics-first (1.0) | Event-first (2.0) | |
|---|---|---|
| Aggregation point | Write time (counters, histograms, recording rules) | Read time (GROUP BY over raw events) |
| Cost driver | Number of distinct label combinations (time series) | Number and size of events stored |
Effect of adding user_id |
Cardinality explosion: one series per user per metric | One more column. Cost barely changes |
| Questions answerable | Only those anticipated when the metric was defined | Any combination of fields present on the event |
| Loss | Detail is discarded at write time and cannot be recovered | Detail is kept. Volume is controlled by sampling instead |
Two properties do most of the work:
- High cardinality — a field with many unique values (
user_id,trace_id,build_sha). These are precisely the fields that locate a specific customer, deploy or shard during an incident. Time-series databases key storage on label combinations, so every new high-cardinality label multiplies series count. - High dimensionality — many fields per event. Each extra field is another axis you can slice by later without re-instrumenting. Columnar stores only read the columns a query touches, so an event with 200 fields costs little more to query than one with 20.
Majors' summary of the consequence: 1.0 tools "can only give you aggregates and random exemplars", while 2.0 tools can tell you precisely what happened when you flipped a flag or deployed a canary (charity.wtf, 2024).
Pre-aggregation does not disappear
Removing metrics as first-class instrumentation does not remove pre-aggregation. It moves it from the application (and the metric SDK) into the database, as materialized views or continuous aggregation computed from raw events.
Wide Event Anatomy¶
A wide event captures the complete context of one unit of work (usually one request in one service) in one structured record. This example from a checkout service shows about 30 fields across 7 context groups:
{
"timestamp": "2025-01-15T10:23:45.612Z",
"request_id": "req_8bf7ec2d",
"trace_id": "abc123",
"service": "checkout-service",
"version": "2.4.1",
"deployment_id": "deploy_789",
"region": "us-east-1",
"method": "POST",
"path": "/api/checkout",
"status_code": 500,
"duration_ms": 1247,
"user": {
"id": "user_456",
"subscription": "premium",
"account_age_days": 847,
"lifetime_value_cents": 284700
},
"cart": {
"id": "cart_xyz",
"item_count": 3,
"total_cents": 15999,
"coupon_applied": "SAVE20"
},
"payment": {
"method": "card",
"provider": "stripe",
"latency_ms": 1089,
"attempt": 3
},
"error": {
"type": "PaymentError",
"code": "card_declined",
"message": "Card declined by issuer",
"retriable": false,
"stripe_decline_code": "insufficient_funds"
},
"feature_flags": {
"new_checkout_flow": true,
"express_payment": false
}
}
When a user complains, filtering on user.id = "user_456" immediately shows:
- Premium customer with a 2+ year account (high priority)
- Payment failed on the 3rd attempt with
insufficient_funds - The new checkout flow was enabled (a correlation worth checking across all failures)
- No grep, no guessing, no second search
The field groups (identity, infrastructure, request, user/business, operation, error, experiments) are listed in Reference. Example queries against events like this are in How-to Guides.
Structured logging is necessary but not sufficient
Structured logging means JSON instead of free text, which is table stakes. Wide events are a discipline: one complete event per unit of work, with all context attached. Structured logs can still be useless: 5 fields, no user context, scattered across 20 lines per request.
Wide Event Lifecycle¶
The implementation idea is simple: create the event when the request starts, let every layer add fields as it learns things, and emit exactly once at the end. The sequence below shows that lifecycle for the checkout example; the runnable code is in How-to Guides.
sequenceDiagram
participant Client
participant Middleware as Wide-event middleware
participant Handler as Checkout handler
participant Deps as Cart DB / payment provider
participant Emitter as Logger or OTel span
participant Backend as Event store
Client->>Middleware: POST /api/checkout
Middleware->>Middleware: Create event (request_id, method, path, service, version)
Middleware->>Handler: next()
Handler->>Handler: Add user.id, user.subscription, lifetime value
Handler->>Deps: getCart(user.id)
Deps-->>Handler: cart
Handler->>Handler: Add cart.id, item_count, total_cents
Handler->>Deps: processPayment(cart, user)
Deps-->>Handler: payment result or decline
Handler->>Handler: Add payment.provider, latency_ms, attempt, error.*
Handler-->>Middleware: response
Middleware->>Middleware: Finalize duration_ms, status_code, outcome
Middleware->>Emitter: Emit one wide event
Emitter->>Backend: One record, about 30 to 50 fields
Middleware-->>Client: HTTP response
Wide Events and OpenTelemetry¶
OpenTelemetry (OTel) is how most teams will transport wide events, but OTel's own model and vocabulary differ from the Observability 2.0 framing. Keeping the terms apart avoids confusion (see also OpenTelemetry and OTel Collector).
OTel's Stance: Separate Signals, Shared Context¶
OTel defines distinct signals (traces, metrics, logs, and the newer profiles signal) and makes them correlatable through shared context (trace and span IDs) and shared resource attributes. An OTel maintainer post states the design directly: "all of your telemetry signals can be correlated through context" and "OpenTelemetry is built on the concept that all signals are interpreted together, rather than separately" (OpenTelemetry Logging and You, 2025-04-18). OTel does not use the "Observability 2.0" label and does not prescribe a single storage model. That choice is left to backends.
Terminology Clash: "Wide Event" vs OTel "Event"¶
| Term | Meaning | Has duration? |
|---|---|---|
| Wide event (Observability 2.0) | One arbitrarily-wide record per unit of work, typically per request per service. In OTel terms it is usually a span enriched with many attributes | Yes (duration_ms) |
| OTel Event | A LogRecord with a mandatory event_name and a schema defined by semantic conventions, emitted via the Logs API. "Not all logs are events, but all events are logs" |
No |
| Span event | A timestamped annotation on a span (Span.AddEvent, Span.RecordException) |
No |
The same OTel post notes that a span is essentially "an event with a particularly detailed schema", distinguished by having a duration and a parent-child hierarchy. That is why Honeycomb-style practice treats enriched spans as the wide events: one span per service hop carries the business context, and the trace links the hops.
Recent OTel Changes That Help Wide Events¶
- Complex attribute types on all signals. Following OTEP 4485 and OTLP 1.9.0, OTel APIs and SDKs are adding maps, heterogeneous arrays, byte arrays and empty values as attribute values on every signal, not only logs. The OTel authors (including Honeycomb's Austin Parker) caution that many backends cannot index or aggregate complex attributes, and advise primitive values where possible (OTel blog, 2025-11-05). In practice, flattened dotted keys (
app.user.id) remain the portable way to build wide spans. - Span Event API deprecation. OTel is deprecating
Span.AddEventandSpan.RecordExceptionin favor of log-based events correlated with the active span; existing span-event data and views keep working through a compatibility layer (OTel blog, 2026-03-17, OTEP 4430). The Events API/SDK had already been marked deprecated in favor of the Logs API withevent_name. For wide-event practice this reinforces the rule: put request context on the span attributes, not in span events.
OpenTelemetry alone does not make telemetry wide
OTel standardizes how telemetry is produced and shipped. It does not decide what context to capture. Auto-instrumentation yields spans with a name, duration, status and HTTP/DB semantic-convention attributes, but no user, cart, plan or feature-flag context. Without deliberate enrichment you get narrow telemetry in a standard format.
When Teams Bypass OTel¶
The OTel pipeline is not always the cheapest path for very high volumes of wide events. ClickHouse reported that its internal LogHouse platform grew beyond 100 PB of uncompressed data, and that it replaced OTel for ClickHouse's own system logs with a purpose-built exporter (SysEx) that preserves native ClickHouse types and avoids intermediate conversions (ClickHouse, June 2025). This is a vendor case study for one specialized source, not a general recommendation to drop OTel.
Sampling in an Event-First World¶
Storing every raw event is the ideal. At high volume it is not affordable, so 2.0 systems control cost with sampling instead of pre-aggregation.
Head vs Tail Sampling¶
- Head sampling decides at the start of a request (usually by trace ID). It is cheap and stateless, but blind: it drops a random 99% whether the request later fails or not.
- Tail sampling decides after the request (or whole trace) completes, using its outcome: status, latency, customer tier, feature flag. It keeps the rare, interesting events and thins the common, healthy ones.
Naive random sampling is dangerous
Sampling 1% of all traffic at random can drop the one request that explains an outage. Use outcome-based rules and keep 100% of errors and slow requests.
Sample Rates Must Travel With the Event¶
Tail sampling skews the stored data on purpose (all errors, few successes). A backend can only compute correct counts and rates if each kept event records the rate it was sampled at, so a query can weight it (an event kept at 1-in-20 counts as 20). Honeycomb's Refinery proxy and its dynamic sampler work this way (Refinery README). If your pipeline does not propagate sample rates, derived error rates will be wrong.
Where Tail Sampling Runs¶
The diagram shows the standard two-tier OpenTelemetry Collector layout: agents route every span of a trace to the same gateway replica, because the tail_sampling processor can only decide once it holds the whole trace in memory.
sequenceDiagram
participant SDK as App with OTel SDK
participant Agent as Collector agent<br/>(loadbalancing exporter)
participant GW as Collector gateway<br/>(tail_sampling processor)
participant Store as Event store
SDK->>Agent: OTLP spans (enriched with app.* attributes)
Agent->>GW: Route by traceID so one replica sees the whole trace
GW->>GW: Buffer spans for decision_wait (default 30s)
GW->>GW: Evaluate policies (status_code, latency, string_attribute, probabilistic)
alt Policy matched
GW->>Store: Export all spans of the trace
else No policy matched
GW->>GW: Drop trace, remember decision in cache
end
The processor is beta for traces and ships in the contrib and k8s Collector distributions. Its README warns that it is stateful: all spans of a trace must reach the same instance, which is why the agent tier uses the loadbalancing exporter (tail sampling processor README). Memory sizing and placement are covered in OTel Collector tail sampling placement. The keep-rate table is in Reference and the config is in How-to Guides.
Database Requirements for Observability 2.0¶
Wide events are large. A single uncompressed wide event can exceed 2 KB (Greptime, 2025-04-25); at 10,000 requests/second that is about 20 MB/s of raw event data before compression. The store must ingest that continuously and serve both real-time dashboards and ad-hoc exploration.
The diagram shows the logical pipeline such a store implements, independent of vendor.
graph LR
subgraph Ingest
OTLP["OTLP receiver"]
TRANSFORM["Ingest-time transform<br/>(parse, enrich, flatten)"]
end
subgraph Store
COLUMNAR["Columnar parts<br/>(per-column encoding and compression)"]
OBJECT["Object storage<br/>(S3 / GCS / Azure Blob)"]
MATVIEW["Materialized views /<br/>continuous aggregation"]
end
subgraph Query
ROUTINE["Routine queries<br/>(dashboards, alerts, SLOs)"]
EXPLORE["Exploratory queries<br/>(ad-hoc SQL, group-by-anything)"]
PROMQL["PromQL<br/>(backward compatibility)"]
end
OTLP --> TRANSFORM --> COLUMNAR
COLUMNAR --> OBJECT
COLUMNAR --> MATVIEW
COLUMNAR --> EXPLORE
MATVIEW --> ROUTINE
MATVIEW --> PROMQL
style COLUMNAR fill:#2980b9,color:#fff
style MATVIEW fill:#8e44ad,color:#fff
| Requirement | Why | How backends implement it |
|---|---|---|
| Columnar storage | Wide events have 50+ fields; queries touch few. Column pruning and vectorized execution avoid reading the rest | ClickHouse MergeTree parts, GreptimeDB Parquet-based SSTs, Honeycomb Retriever segments |
| Disaggregated compute and storage | Retention grows faster than query load; storage must scale independently | Object storage as the primary tier, with local disk or memory caches for recent data |
| Dynamic schema | New fields appear whenever instrumentation changes; an ALTER TABLE per attribute does not scale |
Auto-created columns (GreptimeDB), the JSON type with per-path subcolumns (ClickHouse, production-ready since 25.3), columns created on first appearance of a field (Honeycomb) |
| High-cardinality filtering | user_id, trace_id, request_id have millions of values |
Inverted, skipping and bloom-filter indexes; sort keys on common filters |
| Fresh data | Dashboards and alerts need data within seconds | WAL plus memtable designs; streaming inserts |
| Materialized views | Error rates and p99s must be cheap for dashboards | Incremental aggregation that updates without reprocessing raw events |
| PromQL compatibility | Existing Grafana dashboards and alert rules should keep working | A PromQL engine over the columnar store (GreptimeDB), or metrics tables queried from Grafana |
| Workload isolation | Heavy exploration must not starve alerting | Read replicas or separate compute pools (GreptimeDB Enterprise, ClickHouse Cloud compute-compute separation) |
Routine vs Exploratory Queries¶
| Query type | Purpose | Latency target | Example |
|---|---|---|---|
| Routine | Dashboards, alerts, SLO tracking | Sub-second | Error rate by service over the last 5 minutes |
| Exploratory | Ad-hoc debugging of unknown unknowns | Seconds to minutes | "All requests from user X where flag Y was on and latency > 2 s" |
Routine queries are predictable and can be served from materialized views. Exploratory queries are unpredictable and must scan raw events, which is why they are the workload that separates 2.0 backends from 1.0 ones.
Backend Architectures¶
Three backends illustrate the design space. Their feature facts are tabulated in Reference.
Honeycomb Retriever¶
Honeycomb's SaaS stores all customer events in Retriever, a custom distributed column store. Events arrive through Kafka and are written column-by-column to disk; queries fan out and scan only the columns they need, which is what makes group-by on any field fast. Honeycomb's pitch is the query experience on top: BubbleUp compares a selected slice of events against the baseline to show which dimensions differ (Why Observability Requires a Distributed Column Store). Tail sampling can happen before ingestion in Refinery (Apache-2.0).
ClickHouse and ClickStack¶
ClickStack (launched May 2025) bundles ClickHouse, the HyperDX UI (MIT-licensed, acquired by ClickHouse in March 2025) and an OpenTelemetry Collector configured with an opinionated ClickHouse schema. The project states that "all observability data should be ingested as wide, rich events", but those events are "stored in ClickHouse tables by data type - logs, traces, metrics, and sessions" and correlated at query time (ClickStack README). So ClickStack is event-first storage with per-signal tables, not a single table of wide events. ClickHouse also underpins other OTel-native platforms such as SigNoz.
GreptimeDB Reference Architecture¶
GreptimeDB is an Apache-2.0 (core), Rust-based observability database that runs metrics, logs and traces "on one columnar engine over object storage" with one table model of tags, timestamp and fields (GreptimeDB README). The diagram shows its distributed mode using the project's own component names.
graph TB
subgraph Ingestion
OTLP_IN["OTLP"]
PROM_RW["Prometheus Remote Write"]
OTHER_IN["Loki Push / Elasticsearch Bulk /<br/>InfluxDB line protocol"]
end
subgraph Cluster["GreptimeDB distributed mode"]
FRONTEND["Frontend<br/>(protocol entry, distributed query engine, stateless)"]
METASRV["Metasrv<br/>(metadata, routing, repartitioning)"]
DATANODE["Datanode<br/>(regions: WAL, memtable, SST, indexes, compaction)"]
FLOWNODE["Flownode (optional)<br/>(continuous aggregation / materialized views)"]
KV["etcd or RDS<br/>(metadata KV)"]
end
subgraph Storage
CACHE["Memory and local-disk cache<br/>(recent, hot data)"]
S3["Object storage<br/>(S3 / GCS / Azure Blob)"]
end
subgraph Consumers
GRAFANA["Grafana<br/>(PromQL, SQL)"]
SQL_CLIENT["SQL clients<br/>(MySQL / PostgreSQL wire)"]
JAEGER["Jaeger-compatible trace UI"]
end
OTLP_IN --> FRONTEND
PROM_RW --> FRONTEND
OTHER_IN --> FRONTEND
FRONTEND --> DATANODE
FRONTEND --> METASRV
METASRV --> KV
DATANODE --> CACHE
DATANODE --> S3
DATANODE --> FLOWNODE
FLOWNODE --> FRONTEND
FRONTEND --> GRAFANA
FRONTEND --> SQL_CLIENT
FRONTEND --> JAEGER
style FRONTEND fill:#2980b9,color:#fff
style DATANODE fill:#27ae60,color:#fff
style FLOWNODE fill:#8e44ad,color:#fff
style S3 fill:#e67e22,color:#fff
Key GreptimeDB features for Observability 2.0, per the project README and docs:
- OTLP ingestion — spans land in
opentelemetry_tracesand log records inopentelemetry_logs, both carryingtrace_id, so correlation is a SQL join - Pipelines — ingest-time parsing and transformation
- Flow engine — continuous aggregation (streaming and materialized views) to derive metrics from raw events
- Automatic columns and JSON2 — new attributes become columns; v1.2.0 (2026-09-08) added a structural JSON type stored as structs with dot-style SQL access
- PromQL, SQL and Jaeger-compatible queries, plus MySQL and PostgreSQL wire protocols
- Edition boundary — cluster mode, object storage and Flow are in the Apache-2.0 build; read replicas, workload isolation and automated repartitioning are Enterprise-only
Correction (2026-09)
Earlier versions of this page drew GreptimeDB as "Ingest Nodes", "Query Nodes" and a "Rule Engine". Those are not GreptimeDB component names; the real distributed components are Frontend, Datanode, Metasrv and Flownode. Read replicas are an Enterprise feature, not part of the open-source core.
Common Misconceptions¶
| Misconception | Reality |
|---|---|
| "Structured logging is the same as wide events" | Structured logging is JSON instead of strings. Wide events are a discipline: one event per unit of work with all context attached |
| "We already use OTel, so we're good" | OTel is a delivery mechanism. Auto-instrumented spans carry name, duration, status and protocol attributes. Business context must be added deliberately |
| "This is just tracing with extra steps" | Tracing shows request flow across services. Wide events carry the context within each service. Ideally your wide events are your spans, enriched |
| "Logs are for debugging, metrics are for dashboards" | That split is an artifact of storage engines. Wide events serve both: query them to debug, aggregate them for dashboards |
| "High-cardinality data is expensive and slow" | It is expensive in time-series databases and in logging systems built for full-text search. Columnar stores (ClickHouse, GreptimeDB, Honeycomb Retriever) are built to filter and group on it |
| "2.0 means no metrics" | Even Honeycomb added time-series metrics (GA March 2026). Infrastructure counters such as CPU or queue depth have no request to attach to and remain metrics |
Criticisms and Limits¶
The paradigm is influential but not settled. The strongest objections, including from its proponents:
- The version number is marketing. Hazel Weakly, who proposed an "Observability 3.0" framing, has called the 1.0/2.0/3.0 versioning "entirely marketing", and Majors has said she dislikes the framing while keeping it for its explanatory power (Another observability 3.0 appears on the horizon, 2025-03-24; Hazel Weakly).
- Vendor alignment. The chief advocates sell event stores (Honeycomb, Greptime, ClickHouse). Their cost comparisons are vendor-authored; independent, like-for-like cost studies at scale remain scarce (see the index open questions).
- Not everything is a request. Host metrics, queue depths, garbage-collection counters and batch jobs do not map onto one-event-per-request; they are cheaper as metrics.
- Sampling reintroduces loss. Event-first systems still discard data at high volume. Getting weighting and rules wrong silently skews rates.
- Governance risk. Business context (user IDs, plan tiers, sometimes e-mail or payment data) turns telemetry into a store of personal data. Wide events need attribute allow-lists, redaction and retention rules that metrics never needed.
- Unified storage is often per-signal in practice. ClickStack and GreptimeDB's OTLP ingestion both keep traces and logs in separate tables and correlate them with joins, so "single source of truth" usually means one engine and shared IDs, not literally one table.
Sources¶
- Is It Time To Version Observability? (Signs Point To Yes) — charity.wtf, 2024-08-07
- Live Your Best Life With Structured Events — charity.wtf, 2022-08-15
- How many pillars of observability can you fit on the head of a pin? — charity.wtf, 2025-10-30
- Using Canonical Log Lines for Online Visibility — Brandur Leach, 2016
- OpenTelemetry Logging and You — OTel blog, 2025-04-18
- Announcing Support for Complex Attribute Types in OTel — OTel blog, 2025-11-05
- Deprecating Span Events API — OTel blog, 2026-03-17
- Tail sampling processor README — opentelemetry-collector-contrib
- Observability 2.0 and the Database for It — Greptime, 2025-04-25
- GreptimeDB README
- ClickStack README
- Scaling our Observability platform beyond 100 Petabytes — ClickHouse, 2025
- Refinery README — Honeycomb