Skip to content

Coroot

Summary

Coroot is an open-source (Apache-2.0) observability and APM platform that uses eBPF to collect metrics, logs, traces, and continuous profiles without code changes. It builds a service map automatically, runs built-in inspections against every application, tracks SLOs, fires built-in alerts, and (in Enterprise or via Coroot Cloud) adds AI-assisted root cause analysis. An MCP endpoint lets AI coding agents query the same data. Metrics go to Prometheus (or ClickHouse); logs, traces, and profiles go to ClickHouse. The latest release is v1.26.8 (2026-09-24).

Overview

Coroot is a zero-instrumentation observability platform. A coroot-node-agent DaemonSet attaches eBPF programs to the kernel on every node and captures telemetry for every container, with no SDKs, code changes, or restarts. The product aims to cut Mean Time to Resolution (MTTR): it discovers dependencies automatically, evaluates predefined inspections in each application's context, and sends a single SLO-based alert with the inspection results attached.

Key Facts

Attribute Detail
Repository github.com/coroot/coroot
Stars ~7.9k (2026-09)
Latest Version v1.26.8 (2026-09-24)
Release cadence New minor every 2-5 weeks, patches in between; no LTS or published support window
Operator coroot-operator 1.10.4 (Helm chart 0.10.4, 2026-09-24)
Language Go
License Apache-2.0 (Community Edition, node agent; the agent's BPF code is GPL-2.0); commercial license (Enterprise Edition)
Ownership Vendor-led open source by Coroot (co-founders Nikolay Sivko, Anton Petruhin, Peter Zaitsev); not a CNCF project
Pricing Community free; Enterprise from $1 per CPU core per month
Minimum kernel Linux 5.1 with CONFIG_BPF_EVENTS=y; CO-RE-capable distro for the eBPF profiler
Windows Agent only, Windows Server 2016+ (since v1.23.0)
Storage Prometheus-compatible TSDB or ClickHouse (metrics); ClickHouse (logs, traces, profiles); SQLite or PostgreSQL (configuration)

Evaluation

  • Why it is better: eBPF-based collection covers metrics, logs, traces, and profiles with zero code changes. Service maps, deployment tracking, cost monitoring, and inspections are predefined, so a small team gets useful dashboards on day one. Coroot says its built-in inspections identify "over 80% of issues" automatically (vendor claim, see the README).
  • When it fits (Applicability):
    • Teams that want observability without instrumenting every service
    • Kubernetes, OpenShift, K3s, Docker Swarm, or systemd-based hosts on Linux
    • Organizations that want SLO-driven alerting with root-cause context
    • Small-to-medium teams without a dedicated platform-engineering group
  • When it does not fit: Docker-in-Docker environments such as Minikube, Windows-only estates (the Windows agent sends metrics, logs, and the service map, but no eBPF traces or profiles, and the server needs Linux), or teams that need fully custom Grafana-style dashboards as the primary UI.
Pros Cons
Zero instrumentation required (eBPF) Requires Linux kernel 5.1+ and privileged agents
Auto-generated service maps and deployment tracking Smaller community than Grafana or SigNoz
Predefined inspections and SLO alerts AI RCA is Enterprise-only (or Coroot Cloud for CE, 10 free investigations/month)
Continuous profiling included Less customizable dashboards than Grafana
Apache-2.0 Community Edition eBPF data lacks business-level span context unless you add OTel SDKs
Multi-cluster via remote agents and member projects ClickHouse is required for logs, traces, and profiles
MCP endpoint for AI agents in both editions Fast release pace with no LTS line; plan for frequent upgrades

Architecture

The node agents and the cluster agent push telemetry to the Coroot server, which writes metrics to Prometheus (or ClickHouse) and logs, traces, and profiles to ClickHouse. The server then reads metrics back into its own cache to run inspections.

flowchart LR
    subgraph Node["Each Linux node"]
        NA["coroot-node-agent<br/>(eBPF DaemonSet)"]
    end
    WA["coroot-windows-agent<br/>(ETW, optional)"]
    CA["coroot-cluster-agent<br/>(DB, cloud, K8s events)"]
    SDK["OTel SDK / Collector<br/>(optional)"]
    subgraph Server["Coroot server"]
        Cache["Metric cache<br/>(on-disk)"]
        Insp["Inspections, SLOs,<br/>alerts, RCA, service map"]
        UI["Web UI :8080<br/>MCP endpoint /mcp"]
    end
    Prom["Prometheus / VictoriaMetrics /<br/>Thanos / Mimir"]
    CH["ClickHouse<br/>(logs, traces, profiles,<br/>optional metrics)"]
    NA -->|"Remote Write, OTLP/HTTP,<br/>profiles"| Server
    WA -->|"metrics, logs"| Server
    CA -->|"DB metrics (Remote Write)"| Server
    SDK -->|"OTLP :8080 or :4317"| Server
    Server --> Prom
    Server --> CH
    Prom -->|"PromQL refresh"| Cache
    Cache --> Insp --> UI

Key Components

Component Role
coroot-node-agent eBPF DaemonSet: metrics, logs, traces, and profiles for every container on the node; also runs on Windows (ETW) since v1.23.0
coroot-cluster-agent Database metrics (Postgres, MySQL, Redis/Valkey, Memcached, MongoDB) discovered via the service map; AWS, GCP, and OCI managed-database discovery; Kubernetes events; Go pprof scraping
Coroot server Ingestion endpoints, metric cache, inspections, SLOs, alerting, RCA, UI, MCP endpoint
coroot-operator Kubernetes operator that deploys and auto-upgrades all components from a Coroot custom resource
Storage Prometheus-compatible TSDB or ClickHouse for metrics; ClickHouse for logs, traces, and profiles

The full data-flow, deployment-topology, and trust-boundary diagrams are in Explanation.

Key Features

Feature Detail
eBPF Collection Kernel-level telemetry without code changes; 16 L7 protocols parsed
Service Map Auto-discovery of services, databases, and external APIs
Continuous Profiling eBPF CPU profiler on every node, Go heap profiles and Java async-profiler from the node agent, Go pprof scraping via the cluster agent
Distributed Tracing eBPF-generated spans plus OTLP from OpenTelemetry SDKs
Log Patterns Out-of-the-box event clustering and logs-to-traces correlation (ClickHouse)
SLO Monitoring Availability and latency SLOs with burn-rate incidents
Alerting Rules Built-in and custom rules from checks, log patterns, Kubernetes events, or PromQL (since v1.18.0)
Database Monitoring Deep Postgres, MongoDB, and MySQL inspections (v1.24-v1.26, 2026)
Deployment Tracking Each Kubernetes rollout is compared with the previous one; Flux CD and Argo CD aware
Cost Monitoring Per-application cost for AWS, GCP, and Azure without cloud credentials
AI Root Cause Analysis ML-based dependency walk, then an LLM summary (Enterprise, or Coroot Cloud for CE)
MCP endpoint /mcp exposes topology, health, traces, PromQL, and logs to AI agents; RCA tools are Enterprise-only
Multi-Cluster Agents-only installs report to a central Coroot; memberProjects aggregate several clusters

Editions and Pricing

Tier Cost Adds
Community Free (Apache-2.0) All telemetry, service map, inspections, SLOs, alerting, profiling, multi-cluster, MCP (non-EE tools)
Community + Coroot Cloud 10 free AI investigations per month AI RCA for CE users
Enterprise From $1 per CPU core per month AI RCA with your own LLM, SSO (SAML, OIDC), custom RBAC roles, EE MCP tools, priority support

The full edition matrix is in Reference.

Pricing source

Prices are from the Coroot docs and pricing page, as of 2026-09. Check the pricing page before you budget.

Compatibility

Coroot runs on Kubernetes (EKS, GKE, AKS, OKE, OpenShift, K3s, MicroK8s), Docker, Docker Swarm, and plain systemd hosts with Linux 5.1+. It stores metrics in any Prometheus-compatible TSDB or in ClickHouse, and accepts OTLP over HTTP and gRPC. The full matrix, including CO-RE distributions, Windows, arm64, and unsupported setups, is in Reference.

Recent Developments

  • 2026-09: v1.26.x added deep MySQL monitoring, GCP and OCI managed-database discovery, and service accounts with API keys for MCP and the API (v1.26.8, 2026-09-24).
  • 2026-06 to 2026-08: Windows agent (v1.23.0), native arm64 release images (v1.23.2), deep Postgres (v1.24.0) and MongoDB (v1.25.0) inspections.
  • 2026-04 to 2026-05: Java async-profiler (v1.19.0), Go heap profiling (v1.19.3), and the MCP endpoint (v1.20.x).
  • 2026-02 to 2026-03: built-in alerting rules (v1.18.0), S3 tiering for ClickHouse (v1.18.6), force-SSO (v1.18.7); the legacy coroot Helm chart was dropped in favor of the operator.
  • 2025-10 to 2026-01: ClickHouse as metrics storage (v1.16.0), multi-cluster projects (v1.17.0), remote Coroot data sources (v1.17.7), OIDC SSO (v1.17.9).

The dated release table is in Reference.

Lock-in and Migration

  • Data formats are open. Metrics sit in a Prometheus-compatible TSDB that you can query with PromQL. Traces and logs use the OpenTelemetry ClickHouse layout.
  • Collection is portable. Apps instrumented with OpenTelemetry SDKs keep working if you move to another OTLP backend. eBPF spans, the service map, inspections, and SLO logic are Coroot-specific.
  • Configuration as code. Projects, alerting rules, SLO overrides, integrations, and service accounts can live in the config file or the Coroot resource, which makes rebuilds reproducible.

Community Health

Coroot releases several times a month (23 tagged releases from v1.22.0 on 2026-06-03 to v1.26.8 on 2026-09-24). Support runs through Community Slack, GitHub Discussions, and GitHub issues. The project is company-led rather than foundation-governed.

Topic Map

  • How-to Guides — install with the operator, Docker, Swarm, systemd, or on Windows; projects and API keys; multi-cluster; S3 tiering; TLS; AI and MCP setup; upgrades and troubleshooting
  • Reference — release history, editions and pricing, requirements, ports, flags, CR fields, MCP tools, benchmark results, access tables
  • Explanation — architecture, eBPF data collection, inspections, alerting, RCA, storage model, performance impact, security model

Sources

Source URL
Official website https://coroot.com
Documentation https://docs.coroot.com
Architecture https://docs.coroot.com/installation/architecture/
Requirements https://docs.coroot.com/installation/requirements/
Kubernetes installation https://docs.coroot.com/installation/kubernetes/
Kubernetes operator https://docs.coroot.com/installation/k8s-operator/
Windows agent https://docs.coroot.com/installation/windows/
Configuration https://docs.coroot.com/configuration/configuration/
Performance impact https://docs.coroot.com/installation/performance-impact/
AI root cause analysis https://docs.coroot.com/ai/
Coroot Cloud https://docs.coroot.com/ai/coroot-cloud/
MCP overview https://docs.coroot.com/mcp/overview/
S3 storage for ClickHouse https://docs.coroot.com/guides/clickhouse-s3/
Pricing https://coroot.com/pricing
Enterprise https://coroot.com/enterprise
Repository (server) https://github.com/coroot/coroot
Releases https://github.com/coroot/coroot/releases
Node agent https://github.com/coroot/coroot-node-agent
Cluster agent https://github.com/coroot/coroot-cluster-agent
Operator https://github.com/coroot/coroot-operator
Helm charts https://github.com/coroot/helm-charts
Blog: AI-powered RCA https://coroot.com/blog/we-built-ai-powered-root-cause-analysis-that-actually-works/
Co-founders https://coroot.com/about

Questions

Open Questions

  • What are the real-world limits for Coroot at more than 500 services per cluster? No public benchmark exists; ClickHouse sharding (clickhouse.shards, clickhouse.replicas) and more Coroot replicas backed by PostgreSQL are the documented scaling levers. See Explanation.
  • What do paid Coroot Cloud plans cost beyond the 10 free investigations per month? The docs mention plan details and credit usage but no prices. (The Enterprise trial is 14 days, per coroot.com/enterprise.)
  • Is there a support policy for older minor versions? None is published; the operator's default is to always run the latest release.

Answered Questions

  • What kernel version is required? Linux 5.1 or newer, built with CONFIG_BPF_EVENTS=y. The eBPF profiler also needs CO-RE support. Source: Requirements.
  • Can Coroot work without Kubernetes? Yes. Docker, Docker Swarm, and systemd units are supported. See How-to Guides.
  • Does Coroot run on Windows? The agent does, since v1.23.0 (Windows Server 2016+, x64 and arm64). The server, ClickHouse, and Prometheus need Linux. Source: Windows.
  • What is the eBPF overhead? At 10,000 RPS the latency difference was within measurement error and the agent used about 200m CPU. See Reference.
  • Does Coroot support OpenTelemetry? Yes. It accepts OTLP over HTTP (port 8080) and gRPC (port 4317) for logs and traces.
  • What storage does Coroot use for traces? ClickHouse, with per-signal TTLs (default 7d) applied at table creation. The column-level schema is not documented; see Explanation.
  • Does Coroot run on ARM64 (for example Graviton)? Yes. Release images have been built on native arm64 runners since v1.23.2 (2026-06-30), and the Windows agent ships an arm64 MSI. Source: v1.23.2 release.
  • Which AI models does the Enterprise RCA use? The RCA itself uses ML algorithms that walk the dependency graph. An LLM only summarizes the findings. Supported providers are Anthropic (recommended; Claude Opus 4.6 in the docs), OpenAI (GPT-5.2), and any OpenAI-compatible API. Source: AI configuration.
  • Does Coroot detect Kubernetes NetworkPolicy via eBPF? No. It tracks TCP connections and L7 protocols for the service map and shows the connectivity failures that a bad NetworkPolicy can cause, but it does not read NetworkPolicy objects.