Coroot How-to Guides¶
What this page covers
Task recipes for installing, configuring, scaling, upgrading, and troubleshooting Coroot. Commands follow the official docs as of v1.26.8 (2026-09-24). For flag and field defaults see Reference. For how the parts work see Explanation.
Helm values are Coroot resource fields
The coroot-ce and coroot-ee charts pass their values straight into the spec of a Coroot custom resource. Any --set key must be a CR field, such as agentsOnly.corootURL or clickhouse.s3.mode; see the CR reference. Registry settings (registry.url, registry.pullSecret) belong to the coroot-operator chart instead.
Install¶
This flowchart picks the install path from the officially documented options.
flowchart TD
Start{"Where do the workloads run?"}
Start -->|"Kubernetes or OpenShift"| Multi{"Central Coroot<br/>already running?"}
Multi -->|"No"| Op["coroot-operator +<br/>coroot-ce / coroot-ee chart"]
Multi -->|"Yes"| Agents["coroot-operator +<br/>agentsOnly.corootURL"]
Start -->|"Single Docker host"| Compose["deploy/docker-compose.yaml"]
Start -->|"Docker Swarm"| Swarm["docker-swarm-stack.yaml +<br/>docker run node agent per node"]
Start -->|"Plain Linux VMs"| Systemd["install.sh for server +<br/>node-agent install.sh per host"]
Start -->|"Windows Server"| Win["install.ps1 agent,<br/>server stays on Linux"]
Kubernetes with the operator (recommended)¶
The operator deploys the Coroot server, node agent, cluster agent, kube-state-metrics, Prometheus, and ClickHouse from one Coroot resource, and upgrades them automatically.
helm repo add coroot https://coroot.github.io/helm-charts
helm repo update coroot
helm install -n coroot --create-namespace coroot-operator coroot/coroot-operator
# Community Edition (creates a minimal Coroot custom resource)
helm install -n coroot coroot coroot/coroot-ce \
--set "clickhouse.shards=2,clickhouse.replicas=2"
kubectl port-forward -n coroot service/coroot-coroot 8080:8080
# open http://localhost:8080 and set the admin password
For the Enterprise Edition, install the coroot-ee chart with a license key instead:
helm install -n coroot coroot coroot/coroot-ee \
--set "licenseKey=COROOT-LICENSE-KEY-HERE,clickhouse.shards=2,clickhouse.replicas=2"
Pod Security Standards
If the node agent does not start because of Pod Security violations (common on Talos), allow privileged pods in the namespace: kubectl label ns coroot pod-security.kubernetes.io/enforce=privileged.
Docker Compose¶
This single-host setup runs the Coroot server, node agent, cluster agent, Prometheus, and ClickHouse. Review the compose file first.
curl -fsS https://raw.githubusercontent.com/coroot/coroot/main/deploy/docker-compose.yaml | \
docker compose -f - up -d
docker ps # expect coroot, coroot-node-agent, coroot-cluster-agent, prometheus, clickhouse
# UI: http://localhost:8080 (or http://NODE_IP:8080)
For Enterprise, prefix the docker compose command with LICENSE_KEY="COROOT-LICENSE-KEY-HERE". The compose file then uses the coroot-ee image.
Docker Swarm¶
Deploy the stack on a manager node. Swarm does not support privileged services, so run the node agent on each node by hand.
curl -fsS https://raw.githubusercontent.com/coroot/coroot/main/deploy/docker-swarm-stack.yaml | \
docker stack deploy -c - coroot
# on every node (replace NODE_IP with any Swarm node's IP)
docker run --detach --name coroot-node-agent \
--pull=always --privileged --pid host \
-v /sys/kernel/tracing:/sys/kernel/tracing:rw \
-v /sys/kernel/debug:/sys/kernel/debug:rw \
-v /sys/fs/cgroup:/host/sys/fs/cgroup:ro \
ghcr.io/coroot/coroot-node-agent \
--cgroupfs-root=/host/sys/fs/cgroup \
--collector-endpoint=http://NODE_IP:8080
Linux hosts without containers (systemd)¶
On Ubuntu, Debian, or RHEL you can install everything as systemd services. Install ClickHouse and Prometheus 2.25+ first, and enable the Prometheus remote-write receiver (--enable-feature=remote-write-receiver, or --web.enable-remote-write-receiver on newer releases).
# Coroot server
curl -sfL https://raw.githubusercontent.com/coroot/coroot/main/deploy/install.sh | \
BOOTSTRAP_PROMETHEUS_URL="http://127.0.0.1:9090" \
BOOTSTRAP_REFRESH_INTERVAL=15s \
BOOTSTRAP_CLICKHOUSE_ADDRESS=127.0.0.1:9000 \
sh -
# node agent on every host
curl -sfL https://raw.githubusercontent.com/coroot/coroot-node-agent/main/install.sh | \
COLLECTOR_ENDPOINT=http://COROOT_HOST:8080 \
SCRAPE_INTERVAL=15s \
sh -
Re-run either script to upgrade. Remove the server with /usr/local/bin/coroot-uninstall.sh.
Windows hosts (agent only)¶
Since v1.23.0 the agent runs on Windows Server 2016 or later. Run this in an elevated PowerShell. It installs the coroot-windows-agent service from the MSI package.
$env:COROOT_COLLECTOR_ENDPOINT = 'http://COROOT_URL:8080'
$env:COROOT_API_KEY = '<API_KEY>'
iwr -useb https://raw.githubusercontent.com/coroot/coroot-node-agent/main/install.ps1 | iex
Get-Service coroot-windows-agent
Get-WinEvent -ProviderName coroot-windows-agent -MaxEvents 50 # agent log
Change other options through machine-level COROOT_* environment variables, then run Restart-Service coroot-windows-agent.
Configure¶
Edit the Coroot custom resource¶
All operator settings live in one resource. Edit it in place, or keep a manifest in Git and apply it.
A minimal manifest that uses an existing Prometheus and ClickHouse instead of operator-managed ones:
apiVersion: coroot.com/v1
kind: Coroot
metadata:
name: coroot
namespace: coroot
spec:
externalPrometheus:
url: http://prometheus-server.monitoring:9090 # needs --web.enable-remote-write-receiver
externalClickhouse:
address: clickhouse.observability:9000
database: coroot
user: coroot
passwordSecret:
name: coroot-clickhouse
key: password
To drop Prometheus entirely and keep metrics in ClickHouse, set storeMetricsInClickhouse: true instead of externalPrometheus.
Define projects and API keys in the config file¶
Projects can be created in the UI, or declared in the config file (--config) or the projects list of the Coroot resource. A declared project replaces the API keys and settings of a UI project with the same name.
projects:
- name: production
apiKeys:
- key: ${PRODUCTION_API_KEY} # random string or UUID, must be unique
description: production agents
- name: staging
apiKeys:
- key: ${STAGING_API_KEY}
description: staging agents
In the Coroot resource, use keySecret instead of a plain key. The operator creates the Secret with a random key if it does not exist.
Collect from several clusters¶
Run one full Coroot in a central cluster. In each other cluster, install only the agents and point them at the central instance with that cluster's project API key.
# in each remote cluster
helm install -n coroot --create-namespace coroot-operator coroot/coroot-operator
helm install -n coroot coroot coroot/coroot-ce \
--set "agentsOnly.corootURL=https://coroot.example.com" \
--set "apiKey=REMOTE_PROJECT_API_KEY"
Then add a project that aggregates the per-cluster projects. It holds no API keys and never ingests data:
Tier ClickHouse to S3¶
Since v1.18.6 the operator can place ClickHouse data on S3-compatible storage. Use a dedicated bucket with no external lifecycle rules.
kubectl create secret generic clickhouse-s3-creds -n coroot \
--from-literal=access_key_id=YOUR_ACCESS_KEY \
--from-literal=secret_access_key=YOUR_SECRET_KEY
spec:
clickhouse:
shards: 1
replicas: 2
storage:
size: 100Gi # local disk per replica
s3:
endpoint: https://s3.us-east-1.amazonaws.com/my-coroot-clickhouse/
region: us-east-1
credentials: # omit when using IRSA or workload identity
accessKeyId: {name: clickhouse-s3-creds, key: access_key_id}
secretAccessKey: {name: clickhouse-s3-creds, key: secret_access_key}
cacheSize: 10Gi
mode: tiered # or s3only
moveFactor: "0.1" # move data to S3 when less than 10% of local disk is free
The operator restarts the ClickHouse pods. Check the result from a ClickHouse pod with SELECT name, type FROM system.disks (expect s3_disk and s3_cache).
Enable TLS and an ingress¶
Terminate TLS at an ingress for the UI, or give the Coroot server its own certificate. With tls set, the operator also enables TLS for the OTLP gRPC port.
spec:
ingress:
className: nginx
host: coroot.example.com
tls:
hosts: [coroot.example.com]
secretName: coroot-tls
tls:
certSecret: {name: coroot-server-tls, key: tls.crt}
keySecret: {name: coroot-server-tls, key: tls.key}
Outside Kubernetes, use the server flags --https-listen=:8443 --tls-cert-file=... --tls-key-file=..., and --http-disabled to turn off plain HTTP.
Reduce what the agent collects¶
Use these nodeAgent settings to limit sensitive data or lower overhead:
spec:
nodeAgent:
trackPublicNetworks: ["203.0.113.0/24"] # instead of the default 0.0.0.0/0
logCollector:
collectLogEntries: false # keep log-based metrics, do not store log lines
ebpfTracer:
sampling: "0.1" # keep 10% of eBPF spans
ebpfProfiler:
enabled: true
To exclude a single workload from eBPF profiling, set COROOT_EBPF_PROFILING=disabled in its environment.
Enable Java profiling¶
For better eBPF stacks without async-profiler, start JVMs with -XX:+PreserveFramePointer.
Send OpenTelemetry traces and logs¶
Point OTLP exporters at the Coroot server. HTTP uses port 8080, gRPC uses port 4317.
export OTEL_SERVICE_NAME="checkout" \
OTEL_EXPORTER_OTLP_TRACES_ENDPOINT="http://coroot.coroot:8080/v1/traces" \
OTEL_EXPORTER_OTLP_TRACES_PROTOCOL="http/protobuf" \
OTEL_METRICS_EXPORTER="none"
java -javaagent:opentelemetry-javaagent.jar -jar app.jar
Enable AI Features¶
Root cause analysis in Community Edition (Coroot Cloud)¶
In the UI open Project Settings → Coroot Cloud, create or sign in to an account, and Coroot stores the API key. The free tier gives 10 investigations per month. Turn on Investigate incidents automatically to run RCA on every new incident.
Root cause analysis in Enterprise Edition¶
Open Project Settings → AI (needs settings.edit) and add an Anthropic, OpenAI, or OpenAI-compatible key. As code:
The Coroot server needs outbound HTTPS to the provider, for example api.anthropic.com:443.
Connect an AI agent through MCP¶
For an interactive client, add the server and complete the OAuth flow in the browser:
For a headless agent, create a service account and pass its key as a header:
spec:
serviceAccounts:
- name: investigation-agent
role: Viewer # Viewer can investigate but not resolve alerts
apiKeys:
- description: production investigation agent
keySecret: {name: coroot-investigation-agent, key: api-key}
claude mcp add --transport http coroot https://coroot.example.com/mcp \
--header "Authorization: Bearer $COROOT_API_KEY"
Operate¶
Check health¶
kubectl get coroot -n coroot
kubectl get pods -n coroot -o wide
kubectl logs -n coroot -l app.kubernetes.io/component=coroot-node-agent --tail=50
kubectl logs -n coroot -l app.kubernetes.io/component=coroot-cluster-agent --tail=50
kubectl logs -n coroot -l app.kubernetes.io/component=coroot --tail=50
Useful signals to watch (suggested thresholds, not from Coroot docs):
| Signal | Suggested threshold | Why |
|---|---|---|
| ClickHouse disk usage | Above 70% | The space manager starts dropping the oldest partitions at 70% |
| ClickHouse memory | Above 70% of the limit | Slow queries, OOM kills |
| Node agent restarts | More than 3 per hour | eBPF load failures or memory limits |
| Coroot server data volume | Above 80% | The metric cache fills the storage.size volume |
Upgrade¶
The operator upgrades the Coroot server, node agent, and cluster agent automatically unless you pin image tags in the Coroot resource. Upgrade the operator itself with Helm:
To pin a version, set communityEdition.image.name (or enterpriseEdition.image.name) to a full image reference such as ghcr.io/coroot/coroot:<version>, and change that tag to upgrade. Check the tag format in the registry before you pin.
For Docker Compose, pull and restart:
curl -fsS https://raw.githubusercontent.com/coroot/coroot/main/deploy/docker-compose.yaml | docker compose -f - pull
curl -fsS https://raw.githubusercontent.com/coroot/coroot/main/deploy/docker-compose.yaml | docker compose -f - up -d
Legacy coroot Helm chart
The old all-in-one coroot/coroot chart is deprecated (last app version 1.14.3). Move to coroot-operator plus coroot-ce or coroot-ee.
Uninstall¶
Back up¶
Coroot has no backup command of its own. Back up each store:
- Configuration: the SQLite file in the Coroot data volume (
/data), or the PostgreSQL database. - Metrics: your Prometheus-compatible TSDB's own tooling, for example
vmbackupfor VictoriaMetrics. The Coroot metric cache can be rebuilt from the TSDB. - ClickHouse: a tool such as
clickhouse-backup, or S3 tiering with bucket versioning.
Troubleshoot¶
The node agent collects no data¶
uname -r # must be 5.1 or newer
grep CONFIG_BPF_EVENTS /boot/config-$(uname -r) # must be =y
kubectl get ds -n coroot coroot-node-agent -o jsonpath='{.spec.template.spec.containers[0].securityContext}'
kubectl logs -n coroot -l app.kubernetes.io/component=coroot-node-agent | grep -iE "error|failed|ebpf"
Other common causes: Docker-in-Docker (Minikube) is not supported, the namespace blocks privileged pods, or the agent cannot reach the Coroot URL with a valid API key.
ClickHouse runs out of disk¶
Operator-managed ClickHouse runs as coroot-clickhouse-shard-<n>-<replica> pods and requires the password from the coroot-clickhouse Secret.
kubectl exec -n coroot coroot-clickhouse-shard-0-0 -- sh -c \
'clickhouse-client --password "$CLICKHOUSE_PASSWORD" --query "
SELECT database, table, formatReadableSize(sum(bytes_on_disk)) AS size
FROM system.parts WHERE active GROUP BY database, table ORDER BY sum(bytes_on_disk) DESC LIMIT 10"'
Fixes, in order of preference:
- Increase
clickhouse.storage.size, or add S3 tiering. - Let the space manager work: it drops the oldest partitions above 70% disk usage. Check that
--disable-clickhouse-space-manageris not set. - Lower the TTLs. New TTLs apply only to newly created tables, so change existing tables with
ALTER TABLE ... MODIFY TTLif you need it right away.
For self-managed ClickHouse, Coroot recommends turning off ClickHouse system log tables (query_log, trace_log, metric_log, and similar) with an XML override, because Coroot's heavy query load makes them grow fast.