Coroot Explanation¶
What this page covers
How Coroot collects telemetry with eBPF, where each signal is stored, how inspections, SLOs, alerting, and root cause analysis work, why the performance impact stays low, and the security model. For install and day-2 tasks see How-to Guides. For flags, defaults, pricing, and benchmark numbers see Reference. For key facts see the Coroot hub.
Architecture¶
Components and data flow¶
The Kubernetes deployment below shows the operator-managed layout. Agents only talk to the Coroot server; the server owns all storage connections.
flowchart TB
subgraph K8s["Kubernetes cluster"]
subgraph Agents["Data plane"]
NA["coroot-node-agent<br/>(DaemonSet, eBPF)"]
CA["coroot-cluster-agent<br/>(Deployment)"]
KSM["kube-state-metrics"]
end
subgraph Control["Control plane"]
OP["coroot-operator"]
CR["Coroot CR<br/>(coroot.com/v1)"]
end
subgraph Server["Coroot server"]
Ingest["Ingest endpoints<br/>HTTP :8080, gRPC :4317"]
Cache["Metric cache<br/>(data_dir)"]
Insp["Inspections + SLO engine"]
RCA["RCA engine<br/>(dependency walk)"]
UI["Web UI + /mcp"]
end
subgraph Storage["Storage"]
Prom["Prometheus / VM / Mimir / Thanos"]
CH["ClickHouse + Keeper<br/>(optional S3 tier)"]
CFG["SQLite or PostgreSQL<br/>(configuration)"]
end
end
OTEL["OTel SDK / Collector<br/>(optional)"]
LLM["LLM provider<br/>(Enterprise AI summary)"]
NA -->|"Remote Write, OTLP, profiles"| Ingest
CA -->|"DB metrics (Remote Write)"| Ingest
OTEL -->|"OTLP"| Ingest
Ingest -->|"remote_write"| Prom
Ingest -->|"native protocol"| CH
Prom -->|"PromQL refresh"| Cache
Cache --> Insp --> RCA --> UI
Server --- CFG
RCA -.->|"findings only"| LLM
OP -->|"reconcile"| CR
CR -->|"manages"| Agents
CR -->|"manages"| Server
CR -->|"manages"| Storage
| Signal | Sent by the agent as | Stored in |
|---|---|---|
| Metrics | Prometheus Remote Write (push) or a Prometheus scrape (pull) | Prometheus-compatible TSDB, or ClickHouse with use_clickhouse / storeMetricsInClickhouse (since v1.16.0) |
| Logs | OTLP over HTTP | ClickHouse |
| Traces | OTLP over HTTP | ClickHouse |
| Profiles | Custom HTTP-based protocol | ClickHouse |
| Configuration | n/a | SQLite by default, PostgreSQL via pg_connection_string |
Applications instrumented with OpenTelemetry SDKs, or an OpenTelemetry Collector, send OTLP logs and traces to the same server over HTTP (/v1/traces, /v1/logs on port 8080) or gRPC (port 4317). Coroot merges these spans with the eBPF spans in the same ClickHouse tables, so one trace view covers both.
Coroot keeps its own on-disk metric cache and treats the TSDB as a source that refreshes the cache (every refresh_interval, default 15s). Because of this, Prometheus can run with a short retention, such as a few hours, while Coroot keeps 30 days of cache (cache.ttl, default 30d). The operator-managed Prometheus uses a 2-day retention for this reason. Source: Architecture.
Why storing metrics in ClickHouse is an option
Since v1.16.0 (2025-10-16) Coroot can keep metrics in ClickHouse instead of Prometheus. That leaves one storage system to run, and ClickHouse replication and sharding then cover every signal. With the operator, set storeMetricsInClickhouse: true and it does not install Prometheus.
Deployment topologies¶
In a single cluster, every node runs a node agent that pushes to one Coroot server with operator-managed ClickHouse and Prometheus.
flowchart LR
subgraph Cluster["Kubernetes cluster"]
NA1["node-agent<br/>(node 1)"]
NA2["node-agent<br/>(node 2)"]
NAN["node-agent<br/>(node N)"]
CA["cluster-agent"]
CS["Coroot server"]
CH["ClickHouse<br/>(2 shards x 2 replicas)"]
Prom["Prometheus / VM"]
end
NA1 --> CS
NA2 --> CS
NAN --> CS
CA --> CS
CS --> CH
CS --> Prom
For several clusters, remote clusters run the operator in agentsOnly mode and point corootURL at the central instance. Each remote cluster uses its own project API key. A project with memberProjects can then aggregate several clusters in one view.
flowchart TB
subgraph Central["Central cluster"]
CS["Coroot server<br/>(full install)"]
CH["ClickHouse"]
Prom["Prometheus / VM"]
end
subgraph Remote1["Remote cluster 1"]
NA_R1["node-agents"]
CA_R1["cluster-agent"]
end
subgraph Remote2["Remote cluster 2"]
NA_R2["node-agents"]
CA_R2["cluster-agent"]
end
NA_R1 -->|"agentsOnly.corootURL + apiKey"| CS
CA_R1 --> CS
NA_R2 -->|"agentsOnly.corootURL + apiKey"| CS
CA_R2 --> CS
CS --> CH
CS --> Prom
A project can also set remoteCoroot to use another Coroot instance as its data source, for metrics and ClickHouse access.
Data Collection¶
eBPF-based auto-instrumentation¶
The coroot-node-agent attaches eBPF programs to kernel tracepoints and kprobes, aggregates events in user space, and pushes the results to the Coroot server.
flowchart LR
subgraph Kernel["Linux kernel (5.1+)"]
TP["Tracepoints"]
KP["kprobes / kretprobes"]
UP["uprobes<br/>(TLS libraries)"]
end
subgraph Agent["coroot-node-agent"]
EBPF["eBPF programs"]
Perf["Perf buffer"]
Agg["User-space aggregation,<br/>log pattern extraction"]
WAL["On-disk spool / WAL<br/>(max-spool-size 500MB)"]
end
TP --> EBPF
KP --> EBPF
UP --> EBPF
EBPF --> Perf --> Agg --> WAL
WAL -->|"Remote Write / OTLP / HTTP"| Server["Coroot server"]
The agent buffers data in an on-disk spool when the collector is unreachable (--max-spool-size, default 500MB) and skips containers younger than --min-container-age (default 30s) to limit short-lived job cardinality. Source: node agent flags.
What eBPF captures without code changes¶
| Signal | Kernel attachment point | Data collected |
|---|---|---|
| Network metrics | TCP connect/listen tracepoints and kprobes | Connection latency, failures, retransmits per container pair |
| L7 requests and traces | Socket read/write syscalls, TLS library uprobes | Protocol, method/path or query, status, duration |
| DNS | Parsed as an L7 protocol | Resolution time, failures, queried domains |
| CPU profiling | perf_event sampling (CO-RE) |
On-CPU flame graphs per process |
| Container lifecycle | cgroup and process events | Start/stop, OOM kills, restarts |
Other signals come from outside eBPF. Delay accounting counters (read over Netlink) show CPU and disk wait per container. Cgroup statistics give resource usage. Logs are read from files in /var/log/, journald, and Docker and containerd log files, and the agent extracts log patterns on the node. Cloud metadata services (AWS, GCP, Azure, Hetzner) supply instance type, region, zone, and spot or on-demand status for cost monitoring. Source: node agent README.
The eBPF tracer parses 16 L7 protocols: HTTP, HTTP/2 (gRPC), Postgres, MySQL, Redis, Memcached, MongoDB, Kafka, Cassandra, RabbitMQ, NATS, Dubbo2, DNS, ClickHouse, ZooKeeper, and FoundationDB.
Attachment details
The exact probe names change between agent versions. Treat the attachment column as a guide and read the node agent source for the current list.
Language-level profiling¶
The eBPF profiler captures CPU only. For memory and lock profiles Coroot uses language-specific profilers:
- Go: the node agent reads heap-profile data directly from the memory of Go processes (since v1.19.3). The cluster agent can also scrape pprof endpoints on annotated pods for CPU, block, and mutex profiles.
- Java: with
--enable-java-async-profiler, the node agent loads async-profiler into HotSpot JVMs through the Attach API (since v1.19.0). It collects CPU, allocation, and lock-contention profiles in 60-second sessions, with no JVM flags or restarts. - Symbolization: JVMs started with
-XX:+PreserveFramePointerand Node.js started with--perf-basic-prof-only-functionsexpose perf maps, which improves eBPF stack traces.
Before upload, the agent prunes code paths below 0.25% of the profile total (--profiles-prune-fraction). Source: Profiling overview.
Windows agent¶
Since v1.23.0 (2026-06-24) the same agent binary also runs on Windows Server 2016 and later. It uses Event Tracing for Windows (ETW) instead of eBPF. It discovers Windows services and Docker containers and reports node and container metrics, a service map from TCP connections, DNS metrics, and logs from the Windows Event Log. Only the agent runs on Windows; the Coroot server, ClickHouse, and Prometheus still need Linux. Source: Windows.
Cluster agent discovery¶
The coroot-cluster-agent uses the service map to find databases and then connects to them with credentials that you supply in Coroot:
| Database | Collection method | Example metrics |
|---|---|---|
| PostgreSQL | pg_stat_* views |
Connections, query latency, replication lag |
| MySQL | SHOW STATUS, performance_schema |
Threads, slow queries, buffer pool hit rate |
| Redis / Valkey | INFO |
Memory, clients, hit rate |
| Memcached | stats |
Hit rate, evictions |
| MongoDB | serverStatus |
Operations/sec, locks |
It also discovers AWS RDS and ElastiCache, GCP Cloud SQL and Memorystore, and OCI MySQL HeatWave, PostgreSQL, and OCI Cache. GCP arrived in v1.26.4 and OCI in v1.26.6 (2026-09); AWS discovery can use the pod's IAM role since v1.26.1. It scrapes Go pprof endpoints on pods annotated with coroot.com/profile-scrape and coroot.com/profile-port, and it collects Kubernetes events, which feed event-based alerting rules. Source: Architecture, releases.
Database monitoring grew sharply in 2026: v1.24.0 added deep Postgres inspections, v1.25.0 MongoDB, and v1.26.0 MySQL (InnoDB internals, Galera, Group Replication, and a latency check that ranks the likely cause).
Service Map Generation¶
Coroot builds a real-time service dependency graph from eBPF connection tracking:
- Connection tracking: eBPF programs record every TCP connection (source to destination).
- Container resolution: addresses are mapped to containers and pods through the container runtime and cgroups.
- Service grouping: pods are grouped by Deployment, StatefulSet, DaemonSet, or systemd unit.
- Protocol detection: L7 protocols (HTTP, gRPC, Postgres, MySQL, Redis, Kafka, MongoDB, and the others listed above) are detected from payload patterns.
- Dependency graph: edges carry request rate, latency, and error rate.
Inspections¶
Coroot evaluates every inspection in the application's context, not as a flat alert stream. For example, the Network round-trip time inspection checks latency between an app and the services it depends on. You can override each threshold for one application or a whole project. The application status on the overview page comes from its SLO inspection. Source: Inspections.
| Category | Example inspections |
|---|---|
| SLOs | Availability SLO, latency SLO |
| Instances | Pod restarts, unavailable replicas |
| CPU | Throttling, usage near limits |
| GPU | Utilization, memory |
| Memory | OOM kills, usage near limits |
| Storage | Disk usage, I/O latency |
| Network | Round-trip time, connection errors, retransmits |
| DNS | Latency, errors |
| Logs | Error-rate spikes, new patterns |
| Runtimes | JVM heap and GC, .NET, Python |
| Databases | Postgres, MySQL, MongoDB, Redis, Memcached health |
| Deployments | Rollout tracking, comparison with the previous release |
Inspection count
Earlier notes quoted "18 inspection categories". The docs do not publish a fixed count, so treat the table as representative.
SLO Monitoring¶
Coroot tracks availability and latency SLOs per application from RED metrics (rate, errors, duration):
- Default SLOs are derived from eBPF data, so you do not need to configure them first.
- Error budgets are tracked continuously, and burn-rate breaches open incidents.
- One alert per incident carries the results of all related inspections, which cuts alert noise.
Alerting Rules¶
SLO incidents cover user-facing symptoms. Since v1.18.0 (2026-02-17) Coroot also ships built-in alerting rules for causes such as low disk space or memory pressure. Each rule has one of four source types:
| Source type | What it evaluates |
|---|---|
check |
The result of a built-in inspection check, such as postgres_latency |
log_patterns |
New or frequent log patterns at chosen severities; optional AI evaluation to cut noise |
kubernetes_events |
Kubernetes events collected by the cluster agent; optional AI evaluation |
promql |
Any PromQL expression |
A selector scopes a rule to all applications, to application categories, or to application id patterns. Rules that you define in the config file or the Coroot resource are read-only in the UI. Alerts go to Slack, Microsoft Teams, PagerDuty, Opsgenie, or a webhook, routed per application category. Source: Configuration.
Root Cause Analysis¶
When an SLO incident or anomaly occurs, the RCA runs in two stages. Source: AI-powered RCA.
sequenceDiagram
participant NA as coroot-node-agent
participant CS as Coroot server
participant Insp as Inspections / SLO engine
participant RCA as RCA engine (ML)
participant LLM as LLM provider (EE)
participant Ch as Slack / PagerDuty / Opsgenie / Teams / webhook
NA->>CS: Metrics, traces, logs, profiles
CS->>Insp: Evaluate SLOs and inspections
Insp->>Insp: Burn rate exceeds threshold
Insp->>Ch: Incident alert with inspection results
Insp->>RCA: Investigate anomaly window
RCA->>RCA: Walk dependency graph from the affected service
RCA->>RCA: Test candidate causes (saturation, deploys, DB latency, log spikes, profile shifts)
RCA->>LLM: Send findings only, no raw telemetry
LLM-->>RCA: Summary, propagation paths, suggested fixes
RCA-->>CS: Persist RCA on the incident
- ML stage (no LLM): Coroot follows the dependency graph from the affected service and tests candidate causes against the anomaly, the way an engineer would.
- LLM stage (Enterprise, or Coroot Cloud for Community users): when you click Explain with AI, Coroot sends only its findings to the configured model. The model writes the anomaly summary, propagation paths, and suggested fixes.
Supported providers are Anthropic (recommended by Coroot), OpenAI, and any OpenAI-compatible API such as DeepSeek or Gemini. Source: AI configuration.
Community Edition users get the same flow through Coroot Cloud: after they connect an account, investigations run with 10 free credits per month, and Coroot can investigate new incidents automatically (corootCloud.rca.disableIncidentsAutoInvestigation turns that off). Since v1.19.1 incident notifications carry the RCA summary and remediation steps. Source: Coroot Cloud.
Why an LLM does not do the analysis
Coroot keeps the diagnosis deterministic. The ML stage correlates the affected SLI with every candidate signal and ranks causes. The LLM only turns ranked evidence into prose. Coroot sends findings, not raw telemetry, which limits both token cost and the data that leaves the cluster.
MCP endpoint for AI agents¶
Since v1.20.x (2026-05) Coroot serves a Model Context Protocol endpoint at /mcp, so coding agents such as Claude Code, Cursor, or Codex can investigate production the way an SRE would. Community Edition tools include list_applications, get_application_status, traces_summary, query_metrics, and query_logs. The Enterprise tools list_anomalies and investigate_anomaly run the RCA engine. Interactive clients sign in with OAuth 2.0 and act with the user's permissions. Service accounts with API keys (added in v1.26.8, 2026-09-24) give headless agents programmatic access under the same RBAC rules. Responses are pre-summarized and capped (about 50 KB for lists) to fit agent context limits. The full tool list is in Reference. Source: MCP overview.
Storage Model¶
ClickHouse holds logs, traces, profiles, and optionally metrics:
- Each signal has its own tables. Retention is a table TTL (
traces.ttl,logs.ttl,profiles.ttl,metricsTTL; default7deach). - TTLs apply only at table creation. Changing them later does not alter existing tables.
- The ClickHouse space manager (enabled by default) drops the oldest partitions when disk usage passes 70%, keeping at least one partition per table.
- The operator can tier ClickHouse to S3 (
tieredors3onlymode) since v1.18.6 (2026-03-06). Each shard and replica writes under its own S3 prefix, because ClickHouse zero-copy replication is experimental. With S3 configured, the space manager is disabled and old data moves to S3 instead of being dropped. - Coroot reports compression of about 10x or more for telemetry in ClickHouse (vendor figure, Architecture).
The column-level schema is not documented. Traces and logs follow the OpenTelemetry ClickHouse exporter layout (otel_traces, otel_logs), and profiles use profiling_samples, per Coroot's GitHub issues; read the Coroot source for current definitions. With a shared (global) ClickHouse, Coroot creates a dedicated database per project.
Performance Impact¶
eBPF keeps the agent's cost low for two reasons. First, the kernel verifier checks every eBPF program before it runs: the program must have bounded complexity, so it cannot stall kernel code. Second, events travel to user space through ring buffers. When the agent falls behind, for example under a CPU limit, the kernel drops events instead of blocking the application. Some statistics can then be incomplete, but request latency does not suffer.
Coroot's own benchmark supports this. At a fixed 10,000 RPS against a Go HTTP server, latency with and without coroot-node-agent stayed within measurement error, and the agent used about 200m CPU. Benchmarks of the cluster agent against busy MySQL 8.4, Postgres 18, and MongoDB 8.0 servers (100 databases each) showed no measurable query-latency impact. The database-side cost grows with the number of tables or databases, not with the query rate, because the agent reads statistics views over a single connection. The full test setups and numbers are in Reference.
Vendor numbers
Coroot ran these tests. For workloads well above 10,000 RPS per node, Coroot itself recommends running your own load test. No public benchmark covers more than 500 services per cluster.
CPU profiler overhead¶
| Component | Overhead | Notes |
|---|---|---|
| eBPF CPU profiler | Not published | Coroot's JVM inspection docs describe kernel-level eBPF profiling as "minimal overhead" with no figure (checked 2026-09-28) |
JVM with -XX:+PreserveFramePointer |
About 1-3% | Needed for accurate JVM stacks. Figure from the same JVM inspection docs ("typically around 1-3%") |
Scale caveats¶
- The likely bottlenecks at large scale are the server's metric cache and ClickHouse queries across services. ClickHouse sharding and replication, and more than one Coroot replica with PostgreSQL as the configuration store, are the documented scaling levers.
- eBPF overhead varies with kernel version, workload, and enabled features. If you disable span capture (
ebpfTracer.enabled: false) but keep metrics, the agent uses fewer resources. - Short-lived containers inflate series cardinality. The node agent skips containers younger than 30 seconds by default (
--min-container-age). - Coroot publishes no server or storage sizing tables (the requirements page lists kernel and platform support only, checked 2026-09-28). Size from the vendor benchmarks in Reference and your own load test.
Security Model¶
Components and trust boundaries¶
Agents hold the highest privilege (kernel access) but only reach the Coroot server. The server holds the storage credentials.
flowchart TD
subgraph K8s["Kubernetes cluster"]
subgraph EachNode["Each node"]
Agent["coroot-node-agent<br/>(privileged, hostPID)"]
Workloads["Application pods"]
end
Coroot["Coroot server"]
Operator["coroot-operator"]
end
subgraph Stores["Data stores"]
PG[("SQLite / PostgreSQL<br/>configuration")]
CH[("ClickHouse<br/>traces, logs, profiles")]
Prom[("Prometheus<br/>metrics")]
end
Browser["Browser / API client / MCP agent"]
Workloads -.->|"observed via eBPF"| Agent
Agent -->|"project API key over HTTP(S)"| Coroot
Coroot -->|"write + read"| CH
Coroot -->|"remote_write + PromQL"| Prom
Coroot -->|"config state"| PG
Operator -->|"manages"| Coroot
Operator -->|"manages"| Agent
Browser -->|"session or API key"| Coroot
Node agent privileges¶
The node agent is the most privileged component. It runs as a privileged container in the host PID namespace and mounts the host cgroupfs (read-only), tracefs, and debugfs, because it must attach probes to the kernel and map every process to its container. On clusters that enforce Pod Security Standards, the coroot namespace needs the privileged level. Docker Swarm cannot run privileged services, so there the agent runs as a plain docker run container on each node. The exact settings are in Reference.
Privileged container requirement
The agent runs privileged because bpf() and perf_event_open() need CAP_SYS_ADMIN on kernels older than 5.8. Kernels 5.8+ can use CAP_BPF and CAP_PERFMON instead, but Coroot's manifests still use a privileged container (as of 2026-09).
Per-process opt-out: set COROOT_EBPF_PROFILING=disabled in a workload's environment to exclude it from eBPF profiling.
Data privacy¶
eBPF tracing sees request data at the socket level, so it can capture sensitive information:
- PII in logs: collected container logs can contain user data.
- Database queries: traced SQL can include literal parameter values.
- HTTP headers: request metadata can include cookies or tokens.
| Risk | Mitigation |
|---|---|
| PII in logs | Set logCollector.collectLogEntries: false and keep only log-based metrics |
| Sensitive traces | Lower ebpfTracer.sampling, or disable the tracer; use COROOT_EBPF_PROFILING=disabled on sensitive pods |
| Data accumulation | Keep short TTLs (default 7d) |
| Cross-environment leakage | Use one project with its own API key per environment |
| Unwanted external tracking | Narrow trackPublicNetworks from the default 0.0.0.0/0 |
Authentication and roles¶
Coroot has three identity paths. People log in with a password (or, in Enterprise, through SAML or OIDC SSO, optionally forced). Interactive MCP clients sign in with OAuth 2.0 and inherit the user's role. Headless tools use service accounts, which have no password and authenticate with bearer API keys; Coroot stores only a SHA-256 hash of each key. Anonymous mode removes login entirely and should stay on trusted networks only. The mode table is in Reference.
The built-in roles are Admin, Editor, and Viewer, with fixed permissions. Enterprise adds custom roles that can be limited to one project, one application category, or one application.
Projects are the isolation boundary. Each project has its own apiKeys, which agents use to send telemetry in the X-Api-Key header. A key for production cannot write to staging. A multi-cluster project holds no keys and never ingests data; it reads its member projects on demand. The configuration-file syntax is in How-to Guides.
Network flows¶
All agent traffic flows one way, from the nodes to the Coroot server on port 8080 (or 4317 for OTLP gRPC). Only the server talks to ClickHouse (9000) and Prometheus (9090), and in Enterprise to the LLM provider over HTTPS. So a firewall policy needs only a few rules, and storage credentials never reach the nodes. The port table is in Reference.