Skip to content

OpenObserve

Summary

OpenObserve (O2) is an open-source observability platform written in Rust. It stores logs, metrics, traces, RUM with session replay, continuous profiles, and LLM/agent traces as columnar files (Apache Parquet, optionally Vortex) on object storage, and queries them with SQL (Apache DataFusion) and PromQL. It runs as one binary in single-node mode or as separate Router, Ingester, Querier, Compactor, and Scheduler roles on Kubernetes, backed by PostgreSQL and NATS. The core is AGPL-3.0. A commercial Enterprise edition adds SSO, RBAC, federated search, and audit logs, and is free up to 50 GB/day. v1.0 went GA in September 2026; the latest release is v1.0.4 (2026-09-24).

Overview

OpenObserve, Inc. (formerly Zinc Labs) builds OpenObserve as a lower-cost alternative to Elasticsearch, Splunk, and Datadog. The cost argument has three parts: compressed columnar files on cheap object storage, indexes you turn on only for the fields that need them, and no in-application replication. The vendor says this yields up to 140x lower storage cost than Elasticsearch. That is a vendor claim; the reasoning is in Explanation.

After v1.0 the product covers more than logs, metrics, and traces. It includes RUM and session replay, a service graph, alerts with incidents and SLO burn-rate alerting, synthetic monitoring, database monitoring, continuous profiling over OTLP, and AI observability (agent tracing, evaluations, LLM cost tracking).

Key Facts

Attribute Detail
Repository github.com/openobserve/openobserve
Latest Version v1.0.4 (2026-09-24); v1.0 GA announced 2026-09-22
Previous line v0.92.2 (2026-08-17), the last 0.x release
Release cadence A minor release every 4-7 weeks with patches in between; no LTS or published support window
Language Rust (edition 2024, nightly toolchain); web UI in Vue
Query engine Apache DataFusion 54 (patched fork), Arrow/Parquet 58, Tantivy 0.26 for full-text indexes
License AGPL-3.0 (since November 2023, previously Apache-2.0); Enterprise under a commercial license
Ownership Vendor-led open source by OpenObserve, Inc.; not a foundation project
Pricing OSS free; Self-Hosted Enterprise free up to 50 GB/day; Cloud $0.50/GB ingested + $0.01/GB queried
HA dependencies Kubernetes + Helm, object storage, PostgreSQL, NATS (OpenFGA and Dex for Enterprise RBAC/SSO)
Helm chart openobserve 1.0.2 (app v1.0.1) from https://charts.openobserve.ai
Stars / contributors About 22k stars (2026-09, per the Tools Catalogue); 109 contributors as of 2026-07

Evaluation

  • Why it is better: it is one Rust binary with no JVM and no garbage collector. Object storage keeps storage cheap and durable, and stateless queriers scale quickly. SQL is the main query language (PromQL for metrics). Logs, metrics, traces, RUM, profiles, and LLM traces share one UI and one store. It accepts OTLP, Prometheus remote write, and the Elasticsearch _bulk API, so existing collectors work unchanged.
  • When it fits:
    • Replacing Elasticsearch/ELK for log-heavy workloads where storage cost dominates
    • Teams that want SQL for logs and traces and one tool instead of separate Loki, Tempo, and Mimir
    • Cloud-native environments with S3, GCS, Azure Blob, MinIO, or RustFS
    • Multi-tenant platforms (organizations, per-org ingestion tokens and storage from v0.91)
    • Teams adding LLM/agent observability next to classic telemetry
  • When it does not fit: workloads that need record-level update or delete (data is immutable), heavy wildcard or fuzzy text search without enabling indexes, strict single-vendor-neutral governance (it is not a foundation project), or proprietary SaaS embedding without an AGPL review or a commercial license.
Pros Cons
Rust single binary; no GC pauses AGPL-3.0 core; RBAC/SSO are Enterprise (free only up to 50 GB/day)
Parquet on object storage: cheap, durable, unbounded No in-app replication; unflushed ingester data is a single copy
SQL for logs/traces, SQL or PromQL for metrics Full-text search needs an inverted index or bloom filters per field to be fast
Stateless queriers, compactors, routers, and schedulers HA needs PostgreSQL, NATS, and Kubernetes; not a single binary at scale
RUM, session replay, profiles, synthetics, and LLM observability built in Smaller community and plugin ecosystem than Grafana or Elastic
OTLP, Prometheus remote write, and ES _bulk compatible Fast release pace; docs sometimes lag the source (see Reference)
Multi-tenant organizations; Super Cluster federation (Enterprise) No eBPF auto-instrumentation for traces (profiles can use the OTel eBPF profiler)
SOC 2 Type II and ISO 27001 for the Cloud service Immutable data: no per-record GDPR delete, only retention or stream drop

Architecture

Collectors send data to the Router, which forwards writes to Ingesters and searches to Queriers. Ingesters turn data into Parquet and upload it to object storage. Queriers read it back with DataFusion.

flowchart LR
    SRC["OTel Collector / Prometheus /<br/>Fluent Bit / RUM SDK"] --> RT["Router"]
    RT --> ING["Ingester<br/>(WAL + Memtable)"]
    RT --> QRY["Querier<br/>(DataFusion)"]
    ING -->|Parquet| OBJ[("Object storage")]
    QRY --> OBJ
    CMP["Compactor"] --> OBJ
    SCH["Scheduler<br/>(alerts, reports)"] --> QRY
    ING & QRY & CMP & SCH -.- META[("PostgreSQL<br/>+ NATS")]

In single-node mode one process plays every role, with SQLite and local disk (or object storage). Details are in Explanation.

Component Role
Router Proxies ingest and search requests; serves the UI
Ingester Parses and transforms data, writes the WAL, builds Parquet, uploads to object storage
Querier Runs SQL/PromQL with DataFusion over object storage and ingester buffers
Compactor Merges small files, enforces retention, maintains the file list
Scheduler Runs alerts, reports, and derived streams (formerly alertmanager)

Key Features

Feature Detail
Logs, metrics, traces, profiles One store and UI; profiles via OTLP Profiles
RUM Core Web Vitals, errors, session replay (web and mobile SDKs)
Query languages SQL (DataFusion plus match_all, histogram, approx_topk), PromQL for metrics
Indexes Bloom filters, Tantivy inverted index, KeyValue and hash partitions, all opt-in per field
Dashboards 19+ chart types, drag-and-drop builder, template variables
Alerting Scheduled, real-time, anomaly, composite, and SLO burn-rate alerts; incidents; 1,200+ alert library (v1.0)
Pipelines Visual editor with VRL transforms; routing, redaction, logs-to-metrics
Synthetic monitoring Browser and HTTP checks, private locations (open source from v1.0)
AI observability Agent and session tracing, evaluations, LLM cost and token tracking, agent graph
Security Organizations (all editions); RBAC via OpenFGA and SSO via Dex (Enterprise and Cloud)

Licensing and Pricing

Tier Cost Notes
Open Source $0, AGPL-3.0 Self-hosted; no RBAC or SSO
Self-Hosted Enterprise $0 up to 50 GB/day, then custom SSO, RBAC, audit trail, federated search, redaction, QoS
Cloud (pay as you go) $0.50/GB ingested + $0.01/GB queried 30-day logs/traces, 15-month metrics, unlimited users, 14-day trial

Pricing changed

Earlier versions of this page listed a Cloud "Developer" free tier (200 GB, 15-day retention) and a Pro tier at about $0.60/GB. Those tiers are no longer offered. The current prices and the conflicting README statement about a Cloud free tier are covered in Reference.

Compatibility

Dimension Support
Ingestion OTLP gRPC (5081) and HTTP, Prometheus remote write, Elasticsearch _bulk, JSON API, Fluent Bit, Vector, Filebeat, Kinesis Firehose, GCP Pub/Sub
Query SQL for all signals, PromQL for metrics, REST search API
Object storage S3, GCS, Azure Blob, MinIO, RustFS, OpenStack Swift, Civo, Alibaba OSS
Metadata store SQLite (single node), PostgreSQL (required in cluster mode; MySQL no longer supported)
Deployment Binary (Linux glibc and musl, macOS, Windows), Docker, Helm (HA and standalone charts), OpenObserve Cloud
CPU amd64 and arm64; -simd builds for AVX-512 and NEON

Recent Developments

  • v1.0 GA (September 2026): first 1.x release. It adds AI observability, SLOs, composite alerts, the alert library, Terraform/OpenTofu export, database monitoring, and open-source synthetic monitoring (press release).
  • v0.92.0 (2026-08-07): Synthetic Monitoring and expanded AI observability.
  • v0.91.0 (2026-06-22): Super Org multi-tenancy, org-level ingestion tokens and storage, Tantivy performance work (What's New).
  • Platform changes: the alertmanager role became scheduler, NATS is the only cluster coordinator (etcd removed), MySQL metadata support was removed, and Vortex arrived as an optional storage format.

Alternatives and Migration

  • Closest alternatives: SigNoz (ClickHouse-based, OTel-native), LGTM Stack (Loki, Grafana, Tempo, Mimir on object storage), Victoria Stack (local-SSD, very resource-efficient), Elasticsearch/OpenSearch, and Datadog or Splunk as SaaS.
  • Migration in: point existing Beats, Vector, or Fluent Bit Elasticsearch outputs at /api/{org}/ (_bulk), or switch OTel exporters. Prometheus rules can be imported as alerts with the Enterprise o2 CLI.
  • Lock-in: data sits in open formats (Parquet) in your own bucket. Dashboards, alerts, and pipelines are OpenObserve-specific JSON, exportable through the API, the o2 CLI, or Terraform (alerts and SLOs).

Topic Map

  • How-to Guides: single-node and HA install, object storage and PostgreSQL, ingest logs/metrics/traces, query, tune, secure, config as code, upgrade (v1.0.x).
  • Reference: releases, editions, pricing, node roles, ports, config defaults, API endpoints, SQL functions, index types, hardening checklist.
  • Explanation: node roles, write and query paths, durability, storage formats and indexes, DataFusion engine, pipelines, cost model, security model.

Sources

Source URL
Official website https://openobserve.ai
Documentation https://openobserve.ai/docs/
Architecture https://openobserve.ai/docs/architecture/
HA deployment https://openobserve.ai/docs/administration/deployment/ha-deployment/
Environment variables https://openobserve.ai/docs/environment-variables/
Performance tuning https://openobserve.ai/docs/enterprise-setup/performance/
Pricing https://openobserve.ai/pricing/
Downloads and edition comparison https://openobserve.ai/downloads/
Repository and README https://github.com/openobserve/openobserve
Releases https://github.com/openobserve/openobserve/releases
What's New https://openobserve.ai/whats-new/
Helm charts https://github.com/openobserve/openobserve-helm-chart
O2 CLI https://github.com/openobserve/o2-cli
Kubernetes operator https://github.com/openobserve/o2-k8s-operator
AGPL change rationale https://openobserve.ai/blog/what-are-apache-gpl-and-agpl-licenses-and-why-openobserve-moved-from-apache-to-agpl/
v1.0 GA press release https://www.businesswire.com/news/home/20260922544278/en/OpenObserve-Reaches-v1.0-Bringing-AI-Observability-into-the-Same-Platform-as-Logs-Metrics-Traces-and-RUM
Community Slack https://short.openobserve.ai/community
Blog https://openobserve.ai/blog/

Questions

Open Questions

  • How do the Vortex and Parquet formats compare on storage size and query latency for OpenObserve metrics, and when will Vortex become a recommended default? No published benchmark found as of 2026-09-27.
  • How does DataFusion-based search compare with ClickHouse-based backends (SigNoz, ClickStack) on the same queries? No independent benchmark found as of 2026-09-27.
  • When will the docs catch up with the source defaults (WAL and Memtable sizes, ZO_FEATURE_PER_THREAD_LOCK, MySQL)? The discrepancies are tracked in Reference.

Answered Questions

  • Does the v1.0 line get patch releases after v1.1 ships? Not by policy. SECURITY.md says security fixes go to the current stable release only (pre-releases best effort), and no LTS or older-line support is published (SECURITY.md, checked 2026-09-27).
  • What ingest rate can one node handle? The docs measured about 31 MB/s (about 2.6 TB/day) on an Apple M2 with defaults, and cite 7-30 MB/s per vCPU. OTLP/gRPC gives 60-100% more throughput than HTTP JSON. The largest deployment the vendor cites ingests more than 2 PB/day. The ZO_FEATURE_PER_THREAD_LOCK tuning that older docs recommend is not present in the v1.0.4 source. See Reference.
  • What does AGPL-3.0 mean for SaaS providers embedding OpenObserve? If you modify OpenObserve and offer it over a network, you must make your modified source available to its users. The options are to publish your changes, to keep OpenObserve as an unmodified separate service (whether that avoids a derivative work needs legal review), or to use the Enterprise edition, which ships under a commercial license rather than AGPL. Get legal advice before embedding.
  • How does compaction behave with thousands of streams? Each stream compacts on its own. The compactor loops every 10 s (ZO_COMPACT_INTERVAL) and merges into files up to 2 GB (ZO_COMPACT_MAX_FILE_SIZE=2048). An earlier answer here said 256 MB, which is wrong for current releases. Many streams mean more concurrent merges and more file-list rows in PostgreSQL. No public benchmark covers 10,000+ streams.
  • Can OpenObserve query across buckets or regions? Yes, with Enterprise Super Cluster federated search. The leader cluster fans the query out to the member clusters over gRPC and merges the results. See Explanation.
  • What happens if an ingester crashes mid-flush? The WAL is replayed on restart, which rebuilds Memtables that were not flushed. OpenObserve does not replicate in-flight data between ingesters. If the ingester volume is lost, unflushed data is lost, and with the default ZO_WAL_FSYNC_DISABLED=true a node crash can lose the last unsynced writes. An earlier answer here mentioned a "replication factor > 1"; no such setting is documented. See Explanation.
  • What is the storage format? Apache Parquet with Zstd by default. Vortex can be chosen per stream type (ZO_FILE_FORMAT). Tantivy full-text indexes are stored as .ttv files next to the data files.
  • Is OpenObserve a drop-in Elasticsearch replacement? For ingestion, largely yes: it accepts the _bulk API from Beats, Vector, and Fluent Bit. For queries, no: it uses SQL, not the Elasticsearch Query DSL.
  • How does the 140x cost claim hold up? It is a vendor figure. It combines higher compression, cheaper object storage, and no replica shards. A rough model gives 40-120x. See Explanation.
  • Which query languages are supported? SQL for all signals and PromQL for metrics. See Reference.
  • Can it run as a single binary? Yes. ZO_LOCAL_MODE defaults to true (SQLite plus local disk). See How-to Guides.
  • Is it production-ready? v1.0 went GA in September 2026. The vendor reports thousands of deployments.