Skip to content

Coroot Explanation

What this page covers

How Coroot collects telemetry with eBPF, where each signal is stored, how inspections, SLOs, alerting, and root cause analysis work, why the performance impact stays low, and the security model. For install and day-2 tasks see How-to Guides. For flags, defaults, pricing, and benchmark numbers see Reference. For key facts see the Coroot hub.

Architecture

Components and data flow

The Kubernetes deployment below shows the operator-managed layout. Agents only talk to the Coroot server; the server owns all storage connections.

flowchart TB
    subgraph K8s["Kubernetes cluster"]
        subgraph Agents["Data plane"]
            NA["coroot-node-agent<br/>(DaemonSet, eBPF)"]
            CA["coroot-cluster-agent<br/>(Deployment)"]
            KSM["kube-state-metrics"]
        end
        subgraph Control["Control plane"]
            OP["coroot-operator"]
            CR["Coroot CR<br/>(coroot.com/v1)"]
        end
        subgraph Server["Coroot server"]
            Ingest["Ingest endpoints<br/>HTTP :8080, gRPC :4317"]
            Cache["Metric cache<br/>(data_dir)"]
            Insp["Inspections + SLO engine"]
            RCA["RCA engine<br/>(dependency walk)"]
            UI["Web UI + /mcp"]
        end
        subgraph Storage["Storage"]
            Prom["Prometheus / VM / Mimir / Thanos"]
            CH["ClickHouse + Keeper<br/>(optional S3 tier)"]
            CFG["SQLite or PostgreSQL<br/>(configuration)"]
        end
    end
    OTEL["OTel SDK / Collector<br/>(optional)"]
    LLM["LLM provider<br/>(Enterprise AI summary)"]

    NA -->|"Remote Write, OTLP, profiles"| Ingest
    CA -->|"DB metrics (Remote Write)"| Ingest
    OTEL -->|"OTLP"| Ingest
    Ingest -->|"remote_write"| Prom
    Ingest -->|"native protocol"| CH
    Prom -->|"PromQL refresh"| Cache
    Cache --> Insp --> RCA --> UI
    Server --- CFG
    RCA -.->|"findings only"| LLM
    OP -->|"reconcile"| CR
    CR -->|"manages"| Agents
    CR -->|"manages"| Server
    CR -->|"manages"| Storage
Signal Sent by the agent as Stored in
Metrics Prometheus Remote Write (push) or a Prometheus scrape (pull) Prometheus-compatible TSDB, or ClickHouse with use_clickhouse / storeMetricsInClickhouse (since v1.16.0)
Logs OTLP over HTTP ClickHouse
Traces OTLP over HTTP ClickHouse
Profiles Custom HTTP-based protocol ClickHouse
Configuration n/a SQLite by default, PostgreSQL via pg_connection_string

Applications instrumented with OpenTelemetry SDKs, or an OpenTelemetry Collector, send OTLP logs and traces to the same server over HTTP (/v1/traces, /v1/logs on port 8080) or gRPC (port 4317). Coroot merges these spans with the eBPF spans in the same ClickHouse tables, so one trace view covers both.

Coroot keeps its own on-disk metric cache and treats the TSDB as a source that refreshes the cache (every refresh_interval, default 15s). Because of this, Prometheus can run with a short retention, such as a few hours, while Coroot keeps 30 days of cache (cache.ttl, default 30d). The operator-managed Prometheus uses a 2-day retention for this reason. Source: Architecture.

Why storing metrics in ClickHouse is an option

Since v1.16.0 (2025-10-16) Coroot can keep metrics in ClickHouse instead of Prometheus. That leaves one storage system to run, and ClickHouse replication and sharding then cover every signal. With the operator, set storeMetricsInClickhouse: true and it does not install Prometheus.

Deployment topologies

In a single cluster, every node runs a node agent that pushes to one Coroot server with operator-managed ClickHouse and Prometheus.

flowchart LR
    subgraph Cluster["Kubernetes cluster"]
        NA1["node-agent<br/>(node 1)"]
        NA2["node-agent<br/>(node 2)"]
        NAN["node-agent<br/>(node N)"]
        CA["cluster-agent"]
        CS["Coroot server"]
        CH["ClickHouse<br/>(2 shards x 2 replicas)"]
        Prom["Prometheus / VM"]
    end
    NA1 --> CS
    NA2 --> CS
    NAN --> CS
    CA --> CS
    CS --> CH
    CS --> Prom

For several clusters, remote clusters run the operator in agentsOnly mode and point corootURL at the central instance. Each remote cluster uses its own project API key. A project with memberProjects can then aggregate several clusters in one view.

flowchart TB
    subgraph Central["Central cluster"]
        CS["Coroot server<br/>(full install)"]
        CH["ClickHouse"]
        Prom["Prometheus / VM"]
    end
    subgraph Remote1["Remote cluster 1"]
        NA_R1["node-agents"]
        CA_R1["cluster-agent"]
    end
    subgraph Remote2["Remote cluster 2"]
        NA_R2["node-agents"]
        CA_R2["cluster-agent"]
    end
    NA_R1 -->|"agentsOnly.corootURL + apiKey"| CS
    CA_R1 --> CS
    NA_R2 -->|"agentsOnly.corootURL + apiKey"| CS
    CA_R2 --> CS
    CS --> CH
    CS --> Prom

A project can also set remoteCoroot to use another Coroot instance as its data source, for metrics and ClickHouse access.

Data Collection

eBPF-based auto-instrumentation

The coroot-node-agent attaches eBPF programs to kernel tracepoints and kprobes, aggregates events in user space, and pushes the results to the Coroot server.

flowchart LR
    subgraph Kernel["Linux kernel (5.1+)"]
        TP["Tracepoints"]
        KP["kprobes / kretprobes"]
        UP["uprobes<br/>(TLS libraries)"]
    end
    subgraph Agent["coroot-node-agent"]
        EBPF["eBPF programs"]
        Perf["Perf buffer"]
        Agg["User-space aggregation,<br/>log pattern extraction"]
        WAL["On-disk spool / WAL<br/>(max-spool-size 500MB)"]
    end
    TP --> EBPF
    KP --> EBPF
    UP --> EBPF
    EBPF --> Perf --> Agg --> WAL
    WAL -->|"Remote Write / OTLP / HTTP"| Server["Coroot server"]

The agent buffers data in an on-disk spool when the collector is unreachable (--max-spool-size, default 500MB) and skips containers younger than --min-container-age (default 30s) to limit short-lived job cardinality. Source: node agent flags.

What eBPF captures without code changes

Signal Kernel attachment point Data collected
Network metrics TCP connect/listen tracepoints and kprobes Connection latency, failures, retransmits per container pair
L7 requests and traces Socket read/write syscalls, TLS library uprobes Protocol, method/path or query, status, duration
DNS Parsed as an L7 protocol Resolution time, failures, queried domains
CPU profiling perf_event sampling (CO-RE) On-CPU flame graphs per process
Container lifecycle cgroup and process events Start/stop, OOM kills, restarts

Other signals come from outside eBPF. Delay accounting counters (read over Netlink) show CPU and disk wait per container. Cgroup statistics give resource usage. Logs are read from files in /var/log/, journald, and Docker and containerd log files, and the agent extracts log patterns on the node. Cloud metadata services (AWS, GCP, Azure, Hetzner) supply instance type, region, zone, and spot or on-demand status for cost monitoring. Source: node agent README.

The eBPF tracer parses 16 L7 protocols: HTTP, HTTP/2 (gRPC), Postgres, MySQL, Redis, Memcached, MongoDB, Kafka, Cassandra, RabbitMQ, NATS, Dubbo2, DNS, ClickHouse, ZooKeeper, and FoundationDB.

Attachment details

The exact probe names change between agent versions. Treat the attachment column as a guide and read the node agent source for the current list.

Language-level profiling

The eBPF profiler captures CPU only. For memory and lock profiles Coroot uses language-specific profilers:

  • Go: the node agent reads heap-profile data directly from the memory of Go processes (since v1.19.3). The cluster agent can also scrape pprof endpoints on annotated pods for CPU, block, and mutex profiles.
  • Java: with --enable-java-async-profiler, the node agent loads async-profiler into HotSpot JVMs through the Attach API (since v1.19.0). It collects CPU, allocation, and lock-contention profiles in 60-second sessions, with no JVM flags or restarts.
  • Symbolization: JVMs started with -XX:+PreserveFramePointer and Node.js started with --perf-basic-prof-only-functions expose perf maps, which improves eBPF stack traces.

Before upload, the agent prunes code paths below 0.25% of the profile total (--profiles-prune-fraction). Source: Profiling overview.

Windows agent

Since v1.23.0 (2026-06-24) the same agent binary also runs on Windows Server 2016 and later. It uses Event Tracing for Windows (ETW) instead of eBPF. It discovers Windows services and Docker containers and reports node and container metrics, a service map from TCP connections, DNS metrics, and logs from the Windows Event Log. Only the agent runs on Windows; the Coroot server, ClickHouse, and Prometheus still need Linux. Source: Windows.

Cluster agent discovery

The coroot-cluster-agent uses the service map to find databases and then connects to them with credentials that you supply in Coroot:

Database Collection method Example metrics
PostgreSQL pg_stat_* views Connections, query latency, replication lag
MySQL SHOW STATUS, performance_schema Threads, slow queries, buffer pool hit rate
Redis / Valkey INFO Memory, clients, hit rate
Memcached stats Hit rate, evictions
MongoDB serverStatus Operations/sec, locks

It also discovers AWS RDS and ElastiCache, GCP Cloud SQL and Memorystore, and OCI MySQL HeatWave, PostgreSQL, and OCI Cache. GCP arrived in v1.26.4 and OCI in v1.26.6 (2026-09); AWS discovery can use the pod's IAM role since v1.26.1. It scrapes Go pprof endpoints on pods annotated with coroot.com/profile-scrape and coroot.com/profile-port, and it collects Kubernetes events, which feed event-based alerting rules. Source: Architecture, releases.

Database monitoring grew sharply in 2026: v1.24.0 added deep Postgres inspections, v1.25.0 MongoDB, and v1.26.0 MySQL (InnoDB internals, Galera, Group Replication, and a latency check that ranks the likely cause).

Service Map Generation

Coroot builds a real-time service dependency graph from eBPF connection tracking:

  1. Connection tracking: eBPF programs record every TCP connection (source to destination).
  2. Container resolution: addresses are mapped to containers and pods through the container runtime and cgroups.
  3. Service grouping: pods are grouped by Deployment, StatefulSet, DaemonSet, or systemd unit.
  4. Protocol detection: L7 protocols (HTTP, gRPC, Postgres, MySQL, Redis, Kafka, MongoDB, and the others listed above) are detected from payload patterns.
  5. Dependency graph: edges carry request rate, latency, and error rate.

Inspections

Coroot evaluates every inspection in the application's context, not as a flat alert stream. For example, the Network round-trip time inspection checks latency between an app and the services it depends on. You can override each threshold for one application or a whole project. The application status on the overview page comes from its SLO inspection. Source: Inspections.

Category Example inspections
SLOs Availability SLO, latency SLO
Instances Pod restarts, unavailable replicas
CPU Throttling, usage near limits
GPU Utilization, memory
Memory OOM kills, usage near limits
Storage Disk usage, I/O latency
Network Round-trip time, connection errors, retransmits
DNS Latency, errors
Logs Error-rate spikes, new patterns
Runtimes JVM heap and GC, .NET, Python
Databases Postgres, MySQL, MongoDB, Redis, Memcached health
Deployments Rollout tracking, comparison with the previous release

Inspection count

Earlier notes quoted "18 inspection categories". The docs do not publish a fixed count, so treat the table as representative.

SLO Monitoring

Coroot tracks availability and latency SLOs per application from RED metrics (rate, errors, duration):

  • Default SLOs are derived from eBPF data, so you do not need to configure them first.
  • Error budgets are tracked continuously, and burn-rate breaches open incidents.
  • One alert per incident carries the results of all related inspections, which cuts alert noise.

Alerting Rules

SLO incidents cover user-facing symptoms. Since v1.18.0 (2026-02-17) Coroot also ships built-in alerting rules for causes such as low disk space or memory pressure. Each rule has one of four source types:

Source type What it evaluates
check The result of a built-in inspection check, such as postgres_latency
log_patterns New or frequent log patterns at chosen severities; optional AI evaluation to cut noise
kubernetes_events Kubernetes events collected by the cluster agent; optional AI evaluation
promql Any PromQL expression

A selector scopes a rule to all applications, to application categories, or to application id patterns. Rules that you define in the config file or the Coroot resource are read-only in the UI. Alerts go to Slack, Microsoft Teams, PagerDuty, Opsgenie, or a webhook, routed per application category. Source: Configuration.

Root Cause Analysis

When an SLO incident or anomaly occurs, the RCA runs in two stages. Source: AI-powered RCA.

sequenceDiagram
    participant NA as coroot-node-agent
    participant CS as Coroot server
    participant Insp as Inspections / SLO engine
    participant RCA as RCA engine (ML)
    participant LLM as LLM provider (EE)
    participant Ch as Slack / PagerDuty / Opsgenie / Teams / webhook

    NA->>CS: Metrics, traces, logs, profiles
    CS->>Insp: Evaluate SLOs and inspections
    Insp->>Insp: Burn rate exceeds threshold
    Insp->>Ch: Incident alert with inspection results
    Insp->>RCA: Investigate anomaly window
    RCA->>RCA: Walk dependency graph from the affected service
    RCA->>RCA: Test candidate causes (saturation, deploys, DB latency, log spikes, profile shifts)
    RCA->>LLM: Send findings only, no raw telemetry
    LLM-->>RCA: Summary, propagation paths, suggested fixes
    RCA-->>CS: Persist RCA on the incident
  1. ML stage (no LLM): Coroot follows the dependency graph from the affected service and tests candidate causes against the anomaly, the way an engineer would.
  2. LLM stage (Enterprise, or Coroot Cloud for Community users): when you click Explain with AI, Coroot sends only its findings to the configured model. The model writes the anomaly summary, propagation paths, and suggested fixes.

Supported providers are Anthropic (recommended by Coroot), OpenAI, and any OpenAI-compatible API such as DeepSeek or Gemini. Source: AI configuration.

Community Edition users get the same flow through Coroot Cloud: after they connect an account, investigations run with 10 free credits per month, and Coroot can investigate new incidents automatically (corootCloud.rca.disableIncidentsAutoInvestigation turns that off). Since v1.19.1 incident notifications carry the RCA summary and remediation steps. Source: Coroot Cloud.

Why an LLM does not do the analysis

Coroot keeps the diagnosis deterministic. The ML stage correlates the affected SLI with every candidate signal and ranks causes. The LLM only turns ranked evidence into prose. Coroot sends findings, not raw telemetry, which limits both token cost and the data that leaves the cluster.

MCP endpoint for AI agents

Since v1.20.x (2026-05) Coroot serves a Model Context Protocol endpoint at /mcp, so coding agents such as Claude Code, Cursor, or Codex can investigate production the way an SRE would. Community Edition tools include list_applications, get_application_status, traces_summary, query_metrics, and query_logs. The Enterprise tools list_anomalies and investigate_anomaly run the RCA engine. Interactive clients sign in with OAuth 2.0 and act with the user's permissions. Service accounts with API keys (added in v1.26.8, 2026-09-24) give headless agents programmatic access under the same RBAC rules. Responses are pre-summarized and capped (about 50 KB for lists) to fit agent context limits. The full tool list is in Reference. Source: MCP overview.

Storage Model

ClickHouse holds logs, traces, profiles, and optionally metrics:

  • Each signal has its own tables. Retention is a table TTL (traces.ttl, logs.ttl, profiles.ttl, metricsTTL; default 7d each).
  • TTLs apply only at table creation. Changing them later does not alter existing tables.
  • The ClickHouse space manager (enabled by default) drops the oldest partitions when disk usage passes 70%, keeping at least one partition per table.
  • The operator can tier ClickHouse to S3 (tiered or s3only mode) since v1.18.6 (2026-03-06). Each shard and replica writes under its own S3 prefix, because ClickHouse zero-copy replication is experimental. With S3 configured, the space manager is disabled and old data moves to S3 instead of being dropped.
  • Coroot reports compression of about 10x or more for telemetry in ClickHouse (vendor figure, Architecture).

The column-level schema is not documented. Traces and logs follow the OpenTelemetry ClickHouse exporter layout (otel_traces, otel_logs), and profiles use profiling_samples, per Coroot's GitHub issues; read the Coroot source for current definitions. With a shared (global) ClickHouse, Coroot creates a dedicated database per project.

Performance Impact

eBPF keeps the agent's cost low for two reasons. First, the kernel verifier checks every eBPF program before it runs: the program must have bounded complexity, so it cannot stall kernel code. Second, events travel to user space through ring buffers. When the agent falls behind, for example under a CPU limit, the kernel drops events instead of blocking the application. Some statistics can then be incomplete, but request latency does not suffer.

Coroot's own benchmark supports this. At a fixed 10,000 RPS against a Go HTTP server, latency with and without coroot-node-agent stayed within measurement error, and the agent used about 200m CPU. Benchmarks of the cluster agent against busy MySQL 8.4, Postgres 18, and MongoDB 8.0 servers (100 databases each) showed no measurable query-latency impact. The database-side cost grows with the number of tables or databases, not with the query rate, because the agent reads statistics views over a single connection. The full test setups and numbers are in Reference.

Vendor numbers

Coroot ran these tests. For workloads well above 10,000 RPS per node, Coroot itself recommends running your own load test. No public benchmark covers more than 500 services per cluster.

CPU profiler overhead

Component Overhead Notes
eBPF CPU profiler Not published Coroot's JVM inspection docs describe kernel-level eBPF profiling as "minimal overhead" with no figure (checked 2026-09-28)
JVM with -XX:+PreserveFramePointer About 1-3% Needed for accurate JVM stacks. Figure from the same JVM inspection docs ("typically around 1-3%")

Scale caveats

  • The likely bottlenecks at large scale are the server's metric cache and ClickHouse queries across services. ClickHouse sharding and replication, and more than one Coroot replica with PostgreSQL as the configuration store, are the documented scaling levers.
  • eBPF overhead varies with kernel version, workload, and enabled features. If you disable span capture (ebpfTracer.enabled: false) but keep metrics, the agent uses fewer resources.
  • Short-lived containers inflate series cardinality. The node agent skips containers younger than 30 seconds by default (--min-container-age).
  • Coroot publishes no server or storage sizing tables (the requirements page lists kernel and platform support only, checked 2026-09-28). Size from the vendor benchmarks in Reference and your own load test.

Security Model

Components and trust boundaries

Agents hold the highest privilege (kernel access) but only reach the Coroot server. The server holds the storage credentials.

flowchart TD
    subgraph K8s["Kubernetes cluster"]
        subgraph EachNode["Each node"]
            Agent["coroot-node-agent<br/>(privileged, hostPID)"]
            Workloads["Application pods"]
        end
        Coroot["Coroot server"]
        Operator["coroot-operator"]
    end
    subgraph Stores["Data stores"]
        PG[("SQLite / PostgreSQL<br/>configuration")]
        CH[("ClickHouse<br/>traces, logs, profiles")]
        Prom[("Prometheus<br/>metrics")]
    end
    Browser["Browser / API client / MCP agent"]
    Workloads -.->|"observed via eBPF"| Agent
    Agent -->|"project API key over HTTP(S)"| Coroot
    Coroot -->|"write + read"| CH
    Coroot -->|"remote_write + PromQL"| Prom
    Coroot -->|"config state"| PG
    Operator -->|"manages"| Coroot
    Operator -->|"manages"| Agent
    Browser -->|"session or API key"| Coroot

Node agent privileges

The node agent is the most privileged component. It runs as a privileged container in the host PID namespace and mounts the host cgroupfs (read-only), tracefs, and debugfs, because it must attach probes to the kernel and map every process to its container. On clusters that enforce Pod Security Standards, the coroot namespace needs the privileged level. Docker Swarm cannot run privileged services, so there the agent runs as a plain docker run container on each node. The exact settings are in Reference.

Privileged container requirement

The agent runs privileged because bpf() and perf_event_open() need CAP_SYS_ADMIN on kernels older than 5.8. Kernels 5.8+ can use CAP_BPF and CAP_PERFMON instead, but Coroot's manifests still use a privileged container (as of 2026-09).

Per-process opt-out: set COROOT_EBPF_PROFILING=disabled in a workload's environment to exclude it from eBPF profiling.

Data privacy

eBPF tracing sees request data at the socket level, so it can capture sensitive information:

  • PII in logs: collected container logs can contain user data.
  • Database queries: traced SQL can include literal parameter values.
  • HTTP headers: request metadata can include cookies or tokens.
Risk Mitigation
PII in logs Set logCollector.collectLogEntries: false and keep only log-based metrics
Sensitive traces Lower ebpfTracer.sampling, or disable the tracer; use COROOT_EBPF_PROFILING=disabled on sensitive pods
Data accumulation Keep short TTLs (default 7d)
Cross-environment leakage Use one project with its own API key per environment
Unwanted external tracking Narrow trackPublicNetworks from the default 0.0.0.0/0

Authentication and roles

Coroot has three identity paths. People log in with a password (or, in Enterprise, through SAML or OIDC SSO, optionally forced). Interactive MCP clients sign in with OAuth 2.0 and inherit the user's role. Headless tools use service accounts, which have no password and authenticate with bearer API keys; Coroot stores only a SHA-256 hash of each key. Anonymous mode removes login entirely and should stay on trusted networks only. The mode table is in Reference.

The built-in roles are Admin, Editor, and Viewer, with fixed permissions. Enterprise adds custom roles that can be limited to one project, one application category, or one application.

Projects are the isolation boundary. Each project has its own apiKeys, which agents use to send telemetry in the X-Api-Key header. A key for production cannot write to staging. A multi-cluster project holds no keys and never ingests data; it reads its member projects on demand. The configuration-file syntax is in How-to Guides.

Network flows

All agent traffic flows one way, from the nodes to the Coroot server on port 8080 (or 4317 for OTLP gRPC). Only the server talks to ClickHouse (9000) and Prometheus (9090), and in Enterprise to the LLM provider over HTTPS. So a firewall policy needs only a few rules, and storage credentials never reach the nodes. The port table is in Reference.

Sources