Coroot¶
Summary
Coroot is an open-source (Apache-2.0) observability and APM platform that uses eBPF to collect metrics, logs, traces, and continuous profiles without code changes. It builds a service map automatically, runs built-in inspections against every application, tracks SLOs, fires built-in alerts, and (in Enterprise or via Coroot Cloud) adds AI-assisted root cause analysis. An MCP endpoint lets AI coding agents query the same data. Metrics go to Prometheus (or ClickHouse); logs, traces, and profiles go to ClickHouse. The latest release is v1.26.8 (2026-09-24).
Overview¶
Coroot is a zero-instrumentation observability platform. A coroot-node-agent DaemonSet attaches eBPF programs to the kernel on every node and captures telemetry for every container, with no SDKs, code changes, or restarts. The product aims to cut Mean Time to Resolution (MTTR): it discovers dependencies automatically, evaluates predefined inspections in each application's context, and sends a single SLO-based alert with the inspection results attached.
Key Facts¶
| Attribute | Detail |
|---|---|
| Repository | github.com/coroot/coroot |
| Stars | ~7.9k (2026-09) |
| Latest Version | v1.26.8 (2026-09-24) |
| Release cadence | New minor every 2-5 weeks, patches in between; no LTS or published support window |
| Operator | coroot-operator 1.10.4 (Helm chart 0.10.4, 2026-09-24) |
| Language | Go |
| License | Apache-2.0 (Community Edition, node agent; the agent's BPF code is GPL-2.0); commercial license (Enterprise Edition) |
| Ownership | Vendor-led open source by Coroot (co-founders Nikolay Sivko, Anton Petruhin, Peter Zaitsev); not a CNCF project |
| Pricing | Community free; Enterprise from $1 per CPU core per month |
| Minimum kernel | Linux 5.1 with CONFIG_BPF_EVENTS=y; CO-RE-capable distro for the eBPF profiler |
| Windows | Agent only, Windows Server 2016+ (since v1.23.0) |
| Storage | Prometheus-compatible TSDB or ClickHouse (metrics); ClickHouse (logs, traces, profiles); SQLite or PostgreSQL (configuration) |
Evaluation¶
- Why it is better: eBPF-based collection covers metrics, logs, traces, and profiles with zero code changes. Service maps, deployment tracking, cost monitoring, and inspections are predefined, so a small team gets useful dashboards on day one. Coroot says its built-in inspections identify "over 80% of issues" automatically (vendor claim, see the README).
- When it fits (Applicability):
- Teams that want observability without instrumenting every service
- Kubernetes, OpenShift, K3s, Docker Swarm, or systemd-based hosts on Linux
- Organizations that want SLO-driven alerting with root-cause context
- Small-to-medium teams without a dedicated platform-engineering group
- When it does not fit: Docker-in-Docker environments such as Minikube, Windows-only estates (the Windows agent sends metrics, logs, and the service map, but no eBPF traces or profiles, and the server needs Linux), or teams that need fully custom Grafana-style dashboards as the primary UI.
| Pros | Cons |
|---|---|
| Zero instrumentation required (eBPF) | Requires Linux kernel 5.1+ and privileged agents |
| Auto-generated service maps and deployment tracking | Smaller community than Grafana or SigNoz |
| Predefined inspections and SLO alerts | AI RCA is Enterprise-only (or Coroot Cloud for CE, 10 free investigations/month) |
| Continuous profiling included | Less customizable dashboards than Grafana |
| Apache-2.0 Community Edition | eBPF data lacks business-level span context unless you add OTel SDKs |
| Multi-cluster via remote agents and member projects | ClickHouse is required for logs, traces, and profiles |
| MCP endpoint for AI agents in both editions | Fast release pace with no LTS line; plan for frequent upgrades |
Architecture¶
The node agents and the cluster agent push telemetry to the Coroot server, which writes metrics to Prometheus (or ClickHouse) and logs, traces, and profiles to ClickHouse. The server then reads metrics back into its own cache to run inspections.
flowchart LR
subgraph Node["Each Linux node"]
NA["coroot-node-agent<br/>(eBPF DaemonSet)"]
end
WA["coroot-windows-agent<br/>(ETW, optional)"]
CA["coroot-cluster-agent<br/>(DB, cloud, K8s events)"]
SDK["OTel SDK / Collector<br/>(optional)"]
subgraph Server["Coroot server"]
Cache["Metric cache<br/>(on-disk)"]
Insp["Inspections, SLOs,<br/>alerts, RCA, service map"]
UI["Web UI :8080<br/>MCP endpoint /mcp"]
end
Prom["Prometheus / VictoriaMetrics /<br/>Thanos / Mimir"]
CH["ClickHouse<br/>(logs, traces, profiles,<br/>optional metrics)"]
NA -->|"Remote Write, OTLP/HTTP,<br/>profiles"| Server
WA -->|"metrics, logs"| Server
CA -->|"DB metrics (Remote Write)"| Server
SDK -->|"OTLP :8080 or :4317"| Server
Server --> Prom
Server --> CH
Prom -->|"PromQL refresh"| Cache
Cache --> Insp --> UI
Key Components¶
| Component | Role |
|---|---|
| coroot-node-agent | eBPF DaemonSet: metrics, logs, traces, and profiles for every container on the node; also runs on Windows (ETW) since v1.23.0 |
| coroot-cluster-agent | Database metrics (Postgres, MySQL, Redis/Valkey, Memcached, MongoDB) discovered via the service map; AWS, GCP, and OCI managed-database discovery; Kubernetes events; Go pprof scraping |
| Coroot server | Ingestion endpoints, metric cache, inspections, SLOs, alerting, RCA, UI, MCP endpoint |
| coroot-operator | Kubernetes operator that deploys and auto-upgrades all components from a Coroot custom resource |
| Storage | Prometheus-compatible TSDB or ClickHouse for metrics; ClickHouse for logs, traces, and profiles |
The full data-flow, deployment-topology, and trust-boundary diagrams are in Explanation.
Key Features¶
| Feature | Detail |
|---|---|
| eBPF Collection | Kernel-level telemetry without code changes; 16 L7 protocols parsed |
| Service Map | Auto-discovery of services, databases, and external APIs |
| Continuous Profiling | eBPF CPU profiler on every node, Go heap profiles and Java async-profiler from the node agent, Go pprof scraping via the cluster agent |
| Distributed Tracing | eBPF-generated spans plus OTLP from OpenTelemetry SDKs |
| Log Patterns | Out-of-the-box event clustering and logs-to-traces correlation (ClickHouse) |
| SLO Monitoring | Availability and latency SLOs with burn-rate incidents |
| Alerting Rules | Built-in and custom rules from checks, log patterns, Kubernetes events, or PromQL (since v1.18.0) |
| Database Monitoring | Deep Postgres, MongoDB, and MySQL inspections (v1.24-v1.26, 2026) |
| Deployment Tracking | Each Kubernetes rollout is compared with the previous one; Flux CD and Argo CD aware |
| Cost Monitoring | Per-application cost for AWS, GCP, and Azure without cloud credentials |
| AI Root Cause Analysis | ML-based dependency walk, then an LLM summary (Enterprise, or Coroot Cloud for CE) |
| MCP endpoint | /mcp exposes topology, health, traces, PromQL, and logs to AI agents; RCA tools are Enterprise-only |
| Multi-Cluster | Agents-only installs report to a central Coroot; memberProjects aggregate several clusters |
Editions and Pricing¶
| Tier | Cost | Adds |
|---|---|---|
| Community | Free (Apache-2.0) | All telemetry, service map, inspections, SLOs, alerting, profiling, multi-cluster, MCP (non-EE tools) |
| Community + Coroot Cloud | 10 free AI investigations per month | AI RCA for CE users |
| Enterprise | From $1 per CPU core per month | AI RCA with your own LLM, SSO (SAML, OIDC), custom RBAC roles, EE MCP tools, priority support |
The full edition matrix is in Reference.
Pricing source
Prices are from the Coroot docs and pricing page, as of 2026-09. Check the pricing page before you budget.
Compatibility¶
Coroot runs on Kubernetes (EKS, GKE, AKS, OKE, OpenShift, K3s, MicroK8s), Docker, Docker Swarm, and plain systemd hosts with Linux 5.1+. It stores metrics in any Prometheus-compatible TSDB or in ClickHouse, and accepts OTLP over HTTP and gRPC. The full matrix, including CO-RE distributions, Windows, arm64, and unsupported setups, is in Reference.
Recent Developments¶
- 2026-09: v1.26.x added deep MySQL monitoring, GCP and OCI managed-database discovery, and service accounts with API keys for MCP and the API (v1.26.8, 2026-09-24).
- 2026-06 to 2026-08: Windows agent (v1.23.0), native arm64 release images (v1.23.2), deep Postgres (v1.24.0) and MongoDB (v1.25.0) inspections.
- 2026-04 to 2026-05: Java async-profiler (v1.19.0), Go heap profiling (v1.19.3), and the MCP endpoint (v1.20.x).
- 2026-02 to 2026-03: built-in alerting rules (v1.18.0), S3 tiering for ClickHouse (v1.18.6), force-SSO (v1.18.7); the legacy
corootHelm chart was dropped in favor of the operator. - 2025-10 to 2026-01: ClickHouse as metrics storage (v1.16.0), multi-cluster projects (v1.17.0), remote Coroot data sources (v1.17.7), OIDC SSO (v1.17.9).
The dated release table is in Reference.
Lock-in and Migration¶
- Data formats are open. Metrics sit in a Prometheus-compatible TSDB that you can query with PromQL. Traces and logs use the OpenTelemetry ClickHouse layout.
- Collection is portable. Apps instrumented with OpenTelemetry SDKs keep working if you move to another OTLP backend. eBPF spans, the service map, inspections, and SLO logic are Coroot-specific.
- Configuration as code. Projects, alerting rules, SLO overrides, integrations, and service accounts can live in the config file or the
Corootresource, which makes rebuilds reproducible.
Community Health¶
Coroot releases several times a month (23 tagged releases from v1.22.0 on 2026-06-03 to v1.26.8 on 2026-09-24). Support runs through Community Slack, GitHub Discussions, and GitHub issues. The project is company-led rather than foundation-governed.
Topic Map¶
- How-to Guides — install with the operator, Docker, Swarm, systemd, or on Windows; projects and API keys; multi-cluster; S3 tiering; TLS; AI and MCP setup; upgrades and troubleshooting
- Reference — release history, editions and pricing, requirements, ports, flags, CR fields, MCP tools, benchmark results, access tables
- Explanation — architecture, eBPF data collection, inspections, alerting, RCA, storage model, performance impact, security model
Related Topics¶
- Observability Stacks Comparison — Coroot versus LGTM, Victoria, SigNoz, OpenObserve, SkyWalking and Monoscope
- bpftrace — ad-hoc eBPF tracing on the same kernel hooks
- eBPF Developer Tutorial — how eBPF programs like the node agent's are built
- OpenTelemetry — OTLP senders for app-level spans
- OpenTelemetry Collector — can forward OTLP to Coroot
- SigNoz — another ClickHouse-backed, all-in-one APM
- Victoria Stack — a supported metrics backend
- LGTM Stack — Mimir as a metrics backend
- Kubernetes
- Cilium — another eBPF-based platform on the same nodes
Sources¶
Questions¶
Open Questions¶
- What are the real-world limits for Coroot at more than 500 services per cluster? No public benchmark exists; ClickHouse sharding (
clickhouse.shards,clickhouse.replicas) and more Coroot replicas backed by PostgreSQL are the documented scaling levers. See Explanation. - What do paid Coroot Cloud plans cost beyond the 10 free investigations per month? The docs mention plan details and credit usage but no prices. (The Enterprise trial is 14 days, per coroot.com/enterprise.)
- Is there a support policy for older minor versions? None is published; the operator's default is to always run the latest release.
Answered Questions¶
- What kernel version is required? Linux 5.1 or newer, built with
CONFIG_BPF_EVENTS=y. The eBPF profiler also needs CO-RE support. Source: Requirements. - Can Coroot work without Kubernetes? Yes. Docker, Docker Swarm, and systemd units are supported. See How-to Guides.
- Does Coroot run on Windows? The agent does, since v1.23.0 (Windows Server 2016+, x64 and arm64). The server, ClickHouse, and Prometheus need Linux. Source: Windows.
- What is the eBPF overhead? At 10,000 RPS the latency difference was within measurement error and the agent used about 200m CPU. See Reference.
- Does Coroot support OpenTelemetry? Yes. It accepts OTLP over HTTP (port 8080) and gRPC (port 4317) for logs and traces.
- What storage does Coroot use for traces? ClickHouse, with per-signal TTLs (default 7d) applied at table creation. The column-level schema is not documented; see Explanation.
- Does Coroot run on ARM64 (for example Graviton)? Yes. Release images have been built on native arm64 runners since v1.23.2 (2026-06-30), and the Windows agent ships an arm64 MSI. Source: v1.23.2 release.
- Which AI models does the Enterprise RCA use? The RCA itself uses ML algorithms that walk the dependency graph. An LLM only summarizes the findings. Supported providers are Anthropic (recommended; Claude Opus 4.6 in the docs), OpenAI (GPT-5.2), and any OpenAI-compatible API. Source: AI configuration.
- Does Coroot detect Kubernetes NetworkPolicy via eBPF? No. It tracks TCP connections and L7 protocols for the service map and shows the connectivity failures that a bad NetworkPolicy can cause, but it does not read NetworkPolicy objects.