Skip to content

Coroot How-to Guides

What this page covers

Task recipes for installing, configuring, scaling, upgrading, and troubleshooting Coroot. Commands follow the official docs as of v1.26.8 (2026-09-24). For flag and field defaults see Reference. For how the parts work see Explanation.

Helm values are Coroot resource fields

The coroot-ce and coroot-ee charts pass their values straight into the spec of a Coroot custom resource. Any --set key must be a CR field, such as agentsOnly.corootURL or clickhouse.s3.mode; see the CR reference. Registry settings (registry.url, registry.pullSecret) belong to the coroot-operator chart instead.

Install

This flowchart picks the install path from the officially documented options.

flowchart TD
    Start{"Where do the workloads run?"}
    Start -->|"Kubernetes or OpenShift"| Multi{"Central Coroot<br/>already running?"}
    Multi -->|"No"| Op["coroot-operator +<br/>coroot-ce / coroot-ee chart"]
    Multi -->|"Yes"| Agents["coroot-operator +<br/>agentsOnly.corootURL"]
    Start -->|"Single Docker host"| Compose["deploy/docker-compose.yaml"]
    Start -->|"Docker Swarm"| Swarm["docker-swarm-stack.yaml +<br/>docker run node agent per node"]
    Start -->|"Plain Linux VMs"| Systemd["install.sh for server +<br/>node-agent install.sh per host"]
    Start -->|"Windows Server"| Win["install.ps1 agent,<br/>server stays on Linux"]

The operator deploys the Coroot server, node agent, cluster agent, kube-state-metrics, Prometheus, and ClickHouse from one Coroot resource, and upgrades them automatically.

helm repo add coroot https://coroot.github.io/helm-charts
helm repo update coroot
helm install -n coroot --create-namespace coroot-operator coroot/coroot-operator

# Community Edition (creates a minimal Coroot custom resource)
helm install -n coroot coroot coroot/coroot-ce \
  --set "clickhouse.shards=2,clickhouse.replicas=2"

kubectl port-forward -n coroot service/coroot-coroot 8080:8080
# open http://localhost:8080 and set the admin password

For the Enterprise Edition, install the coroot-ee chart with a license key instead:

helm install -n coroot coroot coroot/coroot-ee \
  --set "licenseKey=COROOT-LICENSE-KEY-HERE,clickhouse.shards=2,clickhouse.replicas=2"

Pod Security Standards

If the node agent does not start because of Pod Security violations (common on Talos), allow privileged pods in the namespace: kubectl label ns coroot pod-security.kubernetes.io/enforce=privileged.

Docker Compose

This single-host setup runs the Coroot server, node agent, cluster agent, Prometheus, and ClickHouse. Review the compose file first.

curl -fsS https://raw.githubusercontent.com/coroot/coroot/main/deploy/docker-compose.yaml | \
  docker compose -f - up -d
docker ps          # expect coroot, coroot-node-agent, coroot-cluster-agent, prometheus, clickhouse
# UI: http://localhost:8080 (or http://NODE_IP:8080)

For Enterprise, prefix the docker compose command with LICENSE_KEY="COROOT-LICENSE-KEY-HERE". The compose file then uses the coroot-ee image.

Docker Swarm

Deploy the stack on a manager node. Swarm does not support privileged services, so run the node agent on each node by hand.

curl -fsS https://raw.githubusercontent.com/coroot/coroot/main/deploy/docker-swarm-stack.yaml | \
  docker stack deploy -c - coroot

# on every node (replace NODE_IP with any Swarm node's IP)
docker run --detach --name coroot-node-agent \
  --pull=always --privileged --pid host \
  -v /sys/kernel/tracing:/sys/kernel/tracing:rw \
  -v /sys/kernel/debug:/sys/kernel/debug:rw \
  -v /sys/fs/cgroup:/host/sys/fs/cgroup:ro \
  ghcr.io/coroot/coroot-node-agent \
  --cgroupfs-root=/host/sys/fs/cgroup \
  --collector-endpoint=http://NODE_IP:8080

Linux hosts without containers (systemd)

On Ubuntu, Debian, or RHEL you can install everything as systemd services. Install ClickHouse and Prometheus 2.25+ first, and enable the Prometheus remote-write receiver (--enable-feature=remote-write-receiver, or --web.enable-remote-write-receiver on newer releases).

# Coroot server
curl -sfL https://raw.githubusercontent.com/coroot/coroot/main/deploy/install.sh | \
  BOOTSTRAP_PROMETHEUS_URL="http://127.0.0.1:9090" \
  BOOTSTRAP_REFRESH_INTERVAL=15s \
  BOOTSTRAP_CLICKHOUSE_ADDRESS=127.0.0.1:9000 \
  sh -

# node agent on every host
curl -sfL https://raw.githubusercontent.com/coroot/coroot-node-agent/main/install.sh | \
  COLLECTOR_ENDPOINT=http://COROOT_HOST:8080 \
  SCRAPE_INTERVAL=15s \
  sh -

Re-run either script to upgrade. Remove the server with /usr/local/bin/coroot-uninstall.sh.

Windows hosts (agent only)

Since v1.23.0 the agent runs on Windows Server 2016 or later. Run this in an elevated PowerShell. It installs the coroot-windows-agent service from the MSI package.

$env:COROOT_COLLECTOR_ENDPOINT = 'http://COROOT_URL:8080'
$env:COROOT_API_KEY = '<API_KEY>'
iwr -useb https://raw.githubusercontent.com/coroot/coroot-node-agent/main/install.ps1 | iex

Get-Service coroot-windows-agent
Get-WinEvent -ProviderName coroot-windows-agent -MaxEvents 50   # agent log

Change other options through machine-level COROOT_* environment variables, then run Restart-Service coroot-windows-agent.

Configure

Edit the Coroot custom resource

All operator settings live in one resource. Edit it in place, or keep a manifest in Git and apply it.

kubectl get coroot -n coroot
kubectl edit coroot coroot -n coroot

A minimal manifest that uses an existing Prometheus and ClickHouse instead of operator-managed ones:

apiVersion: coroot.com/v1
kind: Coroot
metadata:
  name: coroot
  namespace: coroot
spec:
  externalPrometheus:
    url: http://prometheus-server.monitoring:9090   # needs --web.enable-remote-write-receiver
  externalClickhouse:
    address: clickhouse.observability:9000
    database: coroot
    user: coroot
    passwordSecret:
      name: coroot-clickhouse
      key: password

To drop Prometheus entirely and keep metrics in ClickHouse, set storeMetricsInClickhouse: true instead of externalPrometheus.

Define projects and API keys in the config file

Projects can be created in the UI, or declared in the config file (--config) or the projects list of the Coroot resource. A declared project replaces the API keys and settings of a UI project with the same name.

projects:
  - name: production
    apiKeys:
      - key: ${PRODUCTION_API_KEY}      # random string or UUID, must be unique
        description: production agents
  - name: staging
    apiKeys:
      - key: ${STAGING_API_KEY}
        description: staging agents

In the Coroot resource, use keySecret instead of a plain key. The operator creates the Secret with a random key if it does not exist.

Collect from several clusters

Run one full Coroot in a central cluster. In each other cluster, install only the agents and point them at the central instance with that cluster's project API key.

# in each remote cluster
helm install -n coroot --create-namespace coroot-operator coroot/coroot-operator
helm install -n coroot coroot coroot/coroot-ce \
  --set "agentsOnly.corootURL=https://coroot.example.com" \
  --set "apiKey=REMOTE_PROJECT_API_KEY"

Then add a project that aggregates the per-cluster projects. It holds no API keys and never ingests data:

projects:
  - name: prod-global
    memberProjects:
      - prod-eu
      - prod-us

Tier ClickHouse to S3

Since v1.18.6 the operator can place ClickHouse data on S3-compatible storage. Use a dedicated bucket with no external lifecycle rules.

kubectl create secret generic clickhouse-s3-creds -n coroot \
  --from-literal=access_key_id=YOUR_ACCESS_KEY \
  --from-literal=secret_access_key=YOUR_SECRET_KEY
spec:
  clickhouse:
    shards: 1
    replicas: 2
    storage:
      size: 100Gi              # local disk per replica
    s3:
      endpoint: https://s3.us-east-1.amazonaws.com/my-coroot-clickhouse/
      region: us-east-1
      credentials:             # omit when using IRSA or workload identity
        accessKeyId:     {name: clickhouse-s3-creds, key: access_key_id}
        secretAccessKey: {name: clickhouse-s3-creds, key: secret_access_key}
      cacheSize: 10Gi
      mode: tiered             # or s3only
      moveFactor: "0.1"        # move data to S3 when less than 10% of local disk is free

The operator restarts the ClickHouse pods. Check the result from a ClickHouse pod with SELECT name, type FROM system.disks (expect s3_disk and s3_cache).

Enable TLS and an ingress

Terminate TLS at an ingress for the UI, or give the Coroot server its own certificate. With tls set, the operator also enables TLS for the OTLP gRPC port.

spec:
  ingress:
    className: nginx
    host: coroot.example.com
    tls:
      hosts: [coroot.example.com]
      secretName: coroot-tls
  tls:
    certSecret: {name: coroot-server-tls, key: tls.crt}
    keySecret:  {name: coroot-server-tls, key: tls.key}

Outside Kubernetes, use the server flags --https-listen=:8443 --tls-cert-file=... --tls-key-file=..., and --http-disabled to turn off plain HTTP.

Reduce what the agent collects

Use these nodeAgent settings to limit sensitive data or lower overhead:

spec:
  nodeAgent:
    trackPublicNetworks: ["203.0.113.0/24"]   # instead of the default 0.0.0.0/0
    logCollector:
      collectLogEntries: false                # keep log-based metrics, do not store log lines
    ebpfTracer:
      sampling: "0.1"                         # keep 10% of eBPF spans
    ebpfProfiler:
      enabled: true

To exclude a single workload from eBPF profiling, set COROOT_EBPF_PROFILING=disabled in its environment.

Enable Java profiling

spec:
  nodeAgent:
    env:
      - name: ENABLE_JAVA_ASYNC_PROFILER
        value: "true"

For better eBPF stacks without async-profiler, start JVMs with -XX:+PreserveFramePointer.

Send OpenTelemetry traces and logs

Point OTLP exporters at the Coroot server. HTTP uses port 8080, gRPC uses port 4317.

export OTEL_SERVICE_NAME="checkout" \
  OTEL_EXPORTER_OTLP_TRACES_ENDPOINT="http://coroot.coroot:8080/v1/traces" \
  OTEL_EXPORTER_OTLP_TRACES_PROTOCOL="http/protobuf" \
  OTEL_METRICS_EXPORTER="none"
java -javaagent:opentelemetry-javaagent.jar -jar app.jar

Enable AI Features

Root cause analysis in Community Edition (Coroot Cloud)

In the UI open Project Settings → Coroot Cloud, create or sign in to an account, and Coroot stores the API key. The free tier gives 10 investigations per month. Turn on Investigate incidents automatically to run RCA on every new incident.

Root cause analysis in Enterprise Edition

Open Project Settings → AI (needs settings.edit) and add an Anthropic, OpenAI, or OpenAI-compatible key. As code:

spec:
  ai:
    provider: anthropic
    anthropic:
      apiKeySecret: {name: coroot-ai, key: anthropic-api-key}

The Coroot server needs outbound HTTPS to the provider, for example api.anthropic.com:443.

Connect an AI agent through MCP

For an interactive client, add the server and complete the OAuth flow in the browser:

claude mcp add --transport http coroot https://coroot.example.com/mcp

For a headless agent, create a service account and pass its key as a header:

spec:
  serviceAccounts:
    - name: investigation-agent
      role: Viewer                # Viewer can investigate but not resolve alerts
      apiKeys:
        - description: production investigation agent
          keySecret: {name: coroot-investigation-agent, key: api-key}
claude mcp add --transport http coroot https://coroot.example.com/mcp \
  --header "Authorization: Bearer $COROOT_API_KEY"

Operate

Check health

kubectl get coroot -n coroot
kubectl get pods -n coroot -o wide
kubectl logs -n coroot -l app.kubernetes.io/component=coroot-node-agent --tail=50
kubectl logs -n coroot -l app.kubernetes.io/component=coroot-cluster-agent --tail=50
kubectl logs -n coroot -l app.kubernetes.io/component=coroot --tail=50

Useful signals to watch (suggested thresholds, not from Coroot docs):

Signal Suggested threshold Why
ClickHouse disk usage Above 70% The space manager starts dropping the oldest partitions at 70%
ClickHouse memory Above 70% of the limit Slow queries, OOM kills
Node agent restarts More than 3 per hour eBPF load failures or memory limits
Coroot server data volume Above 80% The metric cache fills the storage.size volume

Upgrade

The operator upgrades the Coroot server, node agent, and cluster agent automatically unless you pin image tags in the Coroot resource. Upgrade the operator itself with Helm:

helm repo update coroot
helm upgrade -n coroot coroot-operator coroot/coroot-operator

To pin a version, set communityEdition.image.name (or enterpriseEdition.image.name) to a full image reference such as ghcr.io/coroot/coroot:<version>, and change that tag to upgrade. Check the tag format in the registry before you pin.

For Docker Compose, pull and restart:

curl -fsS https://raw.githubusercontent.com/coroot/coroot/main/deploy/docker-compose.yaml | docker compose -f - pull
curl -fsS https://raw.githubusercontent.com/coroot/coroot/main/deploy/docker-compose.yaml | docker compose -f - up -d

Legacy coroot Helm chart

The old all-in-one coroot/coroot chart is deprecated (last app version 1.14.3). Move to coroot-operator plus coroot-ce or coroot-ee.

Uninstall

helm uninstall coroot -n coroot
helm uninstall coroot-operator -n coroot

Back up

Coroot has no backup command of its own. Back up each store:

  • Configuration: the SQLite file in the Coroot data volume (/data), or the PostgreSQL database.
  • Metrics: your Prometheus-compatible TSDB's own tooling, for example vmbackup for VictoriaMetrics. The Coroot metric cache can be rebuilt from the TSDB.
  • ClickHouse: a tool such as clickhouse-backup, or S3 tiering with bucket versioning.

Troubleshoot

The node agent collects no data

uname -r                                  # must be 5.1 or newer
grep CONFIG_BPF_EVENTS /boot/config-$(uname -r)   # must be =y
kubectl get ds -n coroot coroot-node-agent -o jsonpath='{.spec.template.spec.containers[0].securityContext}'
kubectl logs -n coroot -l app.kubernetes.io/component=coroot-node-agent | grep -iE "error|failed|ebpf"

Other common causes: Docker-in-Docker (Minikube) is not supported, the namespace blocks privileged pods, or the agent cannot reach the Coroot URL with a valid API key.

ClickHouse runs out of disk

Operator-managed ClickHouse runs as coroot-clickhouse-shard-<n>-<replica> pods and requires the password from the coroot-clickhouse Secret.

kubectl exec -n coroot coroot-clickhouse-shard-0-0 -- sh -c \
  'clickhouse-client --password "$CLICKHOUSE_PASSWORD" --query "
     SELECT database, table, formatReadableSize(sum(bytes_on_disk)) AS size
     FROM system.parts WHERE active GROUP BY database, table ORDER BY sum(bytes_on_disk) DESC LIMIT 10"'

Fixes, in order of preference:

  1. Increase clickhouse.storage.size, or add S3 tiering.
  2. Let the space manager work: it drops the oldest partitions above 70% disk usage. Check that --disable-clickhouse-space-manager is not set.
  3. Lower the TTLs. New TTLs apply only to newly created tables, so change existing tables with ALTER TABLE ... MODIFY TTL if you need it right away.

For self-managed ClickHouse, Coroot recommends turning off ClickHouse system log tables (query_log, trace_log, metric_log, and similar) with an XML override, because Coroot's heavy query load makes them grow fast.

Reset the admin password

coroot set-admin-password

Sources