Skip to content

How-to Guides

What this page covers

Task recipes for running SigNoz v0.143.x: installing it with Foundry or Helm, migrating from the deprecated Compose files, upgrading, retention, scaling, collector customization, security setup, and day-2 commands. Versions, ports and config keys are in Reference. The reasons behind these steps are in Explanation.

Docker Compose files and install.sh are deprecated

Since v0.130.0 the install.sh script and the Compose/Swarm files under deploy/ in the SigNoz repository are no longer maintained or distributed. New installs use Foundry (foundryctl). Recipes that cd signoz/deploy/docker/clickhouse-setup or open the UI on port 3301 are stale. The UI now listens on 8080.

Deployment

Install on Docker with Foundry

Goal: a single-machine SigNoz for evaluation or small production.

Prerequisites: Linux or macOS (Windows: WSL 2 with Docker Engine inside WSL, not Docker Desktop), Docker Engine 20.10+ with Compose v2, at least 4 GB of memory for Docker, and free ports 8080, 4317 and 4318 (Docker install docs).

# 1. Install foundryctl (pin with FOUNDRY_VERSION=v0.3.0 if you need reproducibility)
curl -fsSL https://signoz.io/foundry.sh | bash

# 2. Minimal casting
cat > casting.yaml <<'EOF'
apiVersion: v1alpha1
kind: Installation
metadata:
  name: signoz
spec:
  deployment:
    flavor: compose
    mode: docker
EOF

# 3. Validate, render and start
foundryctl cast -f casting.yaml

# 4. Verify
docker ps     # signoz-signoz-0, ingester, signoz-telemetrystore-clickhouse-0-0, keeper, metastore
curl -fsS http://localhost:8080/api/v1/health

Open http://localhost:8080 and create the first admin user. Send OTLP to localhost:4317 (gRPC) or localhost:4318 (HTTP).

To review manifests before you start anything:

foundryctl gauge -f casting.yaml     # check prerequisites
foundryctl forge -f casting.yaml     # render into ./pours and write casting.yaml.lock
cd pours/deployment && docker compose up -d

Pick component backends in the casting

Foundry defaults to PostgreSQL for the metastore and ClickHouse Keeper for coordination. Set spec.metastore.kind: sqlite for a dependency-free single node, or spec.telemetrykeeper.kind: zookeeper if you already run ZooKeeper. Run foundryctl gen to generate example castings for Swarm, systemd, Kubernetes (Helm or Kustomize), ECS, Coolify, Railway and Render.

Install on Kubernetes with Helm

Goal: SigNoz on an existing cluster with the official chart (chart 0.143.0 installs v0.143.0).

helm repo add signoz https://charts.signoz.io
helm repo update

# Release "signoz" in namespace "signoz". The chart README uses "platform" / "my-release".
helm install signoz signoz/signoz \
  --namespace signoz --create-namespace \
  --version 0.143.0 \
  -f override-values.yaml

kubectl -n signoz get pods
kubectl -n signoz port-forward svc/signoz 8080:8080

The chart README lists Kubernetes 1.16+ and Helm 3.0+ as prerequisites. The chart creates a signoz StatefulSet, a signoz-otel-collector Deployment, a signoz-telemetrystore-migrator job, and ClickHouse through the Altinity operator. It uses SQLite as the metastore by default (postgresql.enabled: false). For node and pod telemetry, also install the signoz/k8s-infra chart and point it at the collector.

Foundry can also render Kubernetes output (mode: kubernetes, flavor: helm or kustomize). See the Foundry examples.

Migrate a legacy Compose or Swarm install to Foundry

Goal: move an install.sh / deploy/ deployment to Foundry and keep its data (migration guide).

  1. Install foundryctl. Back up your existing docker-compose.yaml, because SigNoz no longer distributes it and it is your only rollback path.
  2. Write a casting.yaml that reproduces the legacy layout: metastore.kind: sqlite, telemetrykeeper.kind: zookeeper, the ClickHouse macros copied from your existing config.xml, and patches that point the generated volumes at the old ones:

    apiVersion: v1alpha1
    kind: Installation
    metadata:
      name: signoz
    spec:
      deployment:
        flavor: compose        # or swarm
        mode: docker
      metastore:
        kind: sqlite
      telemetrykeeper:
        kind: zookeeper
      telemetrystore:
        spec:
          config:
            data:
              config-0-0.yaml: |
                macros:
                  replica: "example01-01-1"   # copy from your existing config
                  shard: "01"
      patches:
        - target: "deployment/compose.yaml"
          operations:
            - op: replace
              path: /volumes/signoz-telemetrykeeper-0-data/name
              value: signoz-zookeeper-1
            - op: replace
              path: /volumes/signoz-telemetrystore-0-0-data/name
              value: signoz-clickhouse
            - op: replace
              path: /volumes/signoz-metastore-sqlite-0-data/name
              value: signoz-sqlite
            - op: add
              path: /services/signoz-telemetrykeeper-zookeeper-0/user
              value: root
    
  3. Render with foundryctl forge -f casting.yaml and review pours/deployment/: image tags (Foundry uses latest), volumes, and ClickHouse config (now YAML instead of XML).

  4. Stop the old stack without -v: docker compose down (or docker stack rm signoz). This causes downtime.
  5. Deploy: foundryctl cast -f casting.yaml. Schema migrations run as part of cast.
  6. Verify that the UI on 8080 shows old data, OTLP on 4317/4318 accepts new data, and the ClickHouse and ZooKeeper logs are clean.

Upgrades

Upgrade SigNoz safely

  1. Find the required stops with the Upgrade Path Tool and read each stop's guide.
  2. Back up the metastore (SQLite file or Postgres). Dashboards, alerts, pipelines and LLM pricing rules live there.
  3. Upgrade SigNoz and the collector together. v0.143.0 requires signoz-otel-collector v0.144.11.
  4. Wait for async migrations to finish before you move to the next stop.

Foundry (Docker Compose)

curl -fsSL https://signoz.io/foundry.sh | bash         # newer foundryctl = newer generated config
docker pull signoz/signoz:latest                         # cast does not pull floating tags
docker pull signoz/signoz-otel-collector:latest
foundryctl cast -f casting.yaml

Foundry (Docker Swarm)

curl -fsSL https://signoz.io/foundry.sh | bash
docker stack rm signoz        # Swarm configs are immutable. Named volumes are kept.
foundryctl cast -f casting.yaml

Helm

helm repo update
helm search repo signoz/signoz --versions | head
helm -n signoz upgrade signoz signoz/signoz --version 0.143.0 -f override-values.yaml
kubectl -n signoz get pods -w

systemd (binary)

ARCH=$(uname -m | sed 's/x86_64/amd64/;s/aarch64/arm64/')
curl -fsSL "https://github.com/SigNoz/signoz/releases/download/v0.143.0/signoz_linux_${ARCH}.tar.gz" \
  | sudo tar -xz --strip-components=1 -C /opt/signoz
curl -fsSL "https://github.com/SigNoz/signoz-otel-collector/releases/download/v0.144.11/signoz-otel-collector_linux_${ARCH}.tar.gz" \
  | sudo tar -xz --strip-components=1 -C /opt/ingester
sudo foundryctl cast -f casting.yaml
sudo systemctl restart signoz-signoz.service signoz-ingester.service   # prefix = metadata.name

Commands from the v0.143.0 upgrade guide.

v0.143.0 specifics

  • Every user must sign in again, because sessions moved to opaque tokens.
  • If you set only SIGNOZ_TOKENIZER_JWT_SECRET / SIGNOZ_JWT_SECRET, that secret is now ignored. To stay on JWT, see Keep JWT sessions.
  • If you override the collector's traces processors list (Helm values or Foundry config.data), lists are replaced, not merged. Add the two new processors yourself: see Add the AI observability processors.
  • Queries, saved views and alerts that use input.value, gen_ai.prompt, gen_ai.completion and similar keys must switch to gen_ai.input.messages / gen_ai.output.messages.

After any upgrade, check that Settings shows the new version and that the collector log has no unknown type errors.

Configuration

Configure retention

Self-hosted: go to Settings, Workspace, Retention Controls and set traces, logs and metrics separately. Click Save for each signal. The defaults are 15 days for logs and traces and 30 days for metrics. On SigNoz Cloud, ask support for another tier (logs and traces 15/30/90/180 days or 1 year, metrics 1/3/6/13 months).

Do not change TTLs with manual ALTER TABLE ... MODIFY TTL

A manual TTL change is a ClickHouse mutation that can rewrite parts and fight the schema migrator. It also leaves SigNoz's recorded retention out of sync. Use the UI. New settings apply to newly ingested data only, and expired data cannot be recovered.

To inspect the current TTL:

docker exec -it signoz-telemetrystore-clickhouse-0-0 \
  clickhouse-client --query "SHOW CREATE TABLE signoz_traces.signoz_index_v3" | grep -i ttl

Customize the ingester (collector) config

With Foundry, merge a fragment into the generated ingester.yaml. Maps deep-merge, lists are replaced, and null deletes a key. The example below adds a Prometheus scrape and keeps otlp in the pipeline:

spec:
  ingester:
    spec:
      config:
        data:
          ingester.yaml: |
            receivers:
              prometheus:
                config:
                  scrape_configs:
                    - job_name: my-app
                      static_configs:
                        - targets: ["my-app:9090"]
            service:
              pipelines:
                metrics:
                  receivers: [otlp, prometheus]   # restate otlp or it is dropped

With Helm, set otelCollector.config in your values file. Helm also replaces lists wholesale.

For bursty traffic, add memory_limiter before batch in each pipeline you override. The default config ships without it. Size its limits to the container memory:

processors:
  memory_limiter:
    check_interval: 1s
    limit_percentage: 80
    spike_limit_percentage: 20

Add the AI observability processors to a custom collector config

Needed only if you override the traces pipeline (collector v0.144.11+). SigNoz fills groups and default_pricing.rules over OpAMP, so leave them empty:

processors:
  signozspanmapper:
    groups: []
  signozllmpricing:
    attrs:
      model: gen_ai.request.model
      in: gen_ai.usage.input_tokens
      out: gen_ai.usage.output_tokens
      cache_read: gen_ai.usage.cache_read.input_tokens
      cache_write: gen_ai.usage.cache_creation.input_tokens
    default_pricing:
      rules: []
    output_attrs:
      in: signoz.gen_ai.usage.input_tokens.cost
      out: signoz.gen_ai.usage.output_tokens.cost
      cache_read: signoz.gen_ai.usage.cache_read.input_tokens.cost
      cache_write: signoz.gen_ai.usage.cache_write.input_tokens.cost
      total: signoz.gen_ai.usage.tokens.cost
service:
  pipelines:
    traces:
      processors: [signozspanmetrics/delta, signozspanmapper, signozllmpricing, batch]

Send data to SigNoz Cloud

Point an existing OpenTelemetry Collector at your region's ingest endpoint with a write-only ingestion key (Settings, Ingestion Settings):

exporters:
  otlp:
    endpoint: "ingest.<region>.signoz.cloud:443"
    tls:
      insecure: false
    headers:
      signoz-ingestion-key: "${env:SIGNOZ_INGESTION_KEY}"

Self-hosted SigNoz does not use ingestion keys, and OTLP is unauthenticated there. See Restrict OTLP ingestion.

Enable the SigNoz MCP server (self-hosted)

spec:
  mcp:
    spec:
      enabled: true
foundryctl cast -f casting.yaml
curl -fsS localhost:8000/livez && echo " OK"
# Register with an AI client using a service-account key, for example Claude Code:
claude mcp add --scope user --transport http signoz http://localhost:8000/mcp \
  --header "SIGNOZ-API-KEY: ${SIGNOZ_API_KEY}"

Scaling

Scale ClickHouse on Kubernetes (shards and replicas)

To run 2 shards x 2 replicas with a 3-node ZooKeeper, put this in override-values.yaml (distributed ClickHouse docs):

clickhouse:
  layout:
    shardsCount: 2
    replicasCount: 2
  zookeeper:
    replicaCount: 3
telemetryStoreMigrator:
  enableReplication: true   # older docs call this schemaMigrator.enableReplication
helm -n signoz upgrade signoz signoz/signoz -f override-values.yaml

With Foundry, set spec.telemetrystore.spec.cluster.shards and .replicas, and the telemetrykeeper replicas.

Note

values.yaml marks clickhouse.layout as experimental. Adding shards does not rebalance existing data. Only new parts land on new shards.

Scale the ingester and SigNoz

Component How to scale
Ingester Add replicas behind a load balancer (Helm otelCollector.replicaCount, Foundry spec.ingester.spec.cluster.replicas). For gRPC, use an L7-aware balancer.
SigNoz binary Scale vertically first. Multi-replica behaviour depends on the sharder config (default noop).
ClickHouse Shard for write throughput. Replicate for HA and read capacity.
Metastore Move from SQLite to PostgreSQL (managed DB recommended) for production HA.

Use the sizing tables in Reference.

Security Setup

Create a service account API key

  1. Go to Settings, Service Accounts, New Service Account. Names use lowercase letters, numbers and hyphens, up to 50 characters.
  2. Assign the narrowest role that works, then open Keys, Add Key. Set an expiry date. The key is shown only once.
  3. Test it:
curl -fsS -H "SIGNOZ-API-KEY: ${SIGNOZ_API_KEY}" https://signoz.example.com/api/v1/service_accounts/me

Manage SigNoz as code with Terraform

terraform {
  required_providers {
    signoz = { source = "signoz/signoz" }
  }
}

provider "signoz" {
  endpoint     = "https://signoz.example.com"   # or env SIGNOZ_ENDPOINT
  access_token = var.signoz_access_token        # prefer env SIGNOZ_ACCESS_TOKEN
}

Requires Terraform 1.4+ or OpenTofu 1.6+ (provider README).

Keep JWT sessions

Since v0.143.0 the default is opaque tokens. To keep JWT, set both variables. SigNoz refuses to start with jwt and no secret:

SIGNOZ_TOKENIZER_PROVIDER=jwt
SIGNOZ_TOKENIZER_JWT_SECRET="$(openssl rand -hex 32)"

Configure SSO

  1. Go to Settings, Organization Settings, Authenticated Domains and add your email domain.
  2. Click Configure SSO and pick Google Workspace (all editions), or SAML or OIDC (Cloud or Enterprise Self-Hosted).
  3. For SAML, enter the IdP SSO URL, X.509 certificate and entity ID. Map IdP groups to the VIEWER, EDITOR or ADMIN role.
  4. Test from a private window, then turn on Enforce SSO. If you get locked out, /login?password=Y still offers password login.

Provider guides: Okta, Microsoft Entra ID, Keycloak.

Restrict OTLP ingestion (self-hosted)

  • Do not publish 4317/4318 to the internet. With Foundry, remove the port mappings with a patches entry. With Helm, keep the collector Service as ClusterIP.
  • If agents must cross untrusted networks, put a gateway collector or proxy in front that terminates TLS and checks a header or mTLS. Then forward to the ingester over the private network.

Create a dedicated ClickHouse user

A users.d fragment (XML shown; Foundry-generated ClickHouse config is YAML with the same keys):

<clickhouse>
  <users>
    <signoz>
      <password_sha256_hex>REPLACE_WITH_SHA256</password_sha256_hex>
      <networks><ip>10.0.0.0/8</ip></networks>
      <profile>default</profile>
      <quota>default</quota>
      <access_management>0</access_management>
    </signoz>
  </users>
</clickhouse>
GRANT SELECT, INSERT ON signoz_traces.* TO signoz;
GRANT SELECT, INSERT ON signoz_logs.* TO signoz;
GRANT SELECT, INSERT ON signoz_metrics.* TO signoz;
GRANT SELECT, INSERT ON signoz_meter.* TO signoz;
GRANT SELECT, INSERT ON signoz_metadata.* TO signoz;
GRANT SELECT, INSERT ON signoz_analytics.* TO signoz;

Run migrations as a user that also has CREATE, ALTER and DROP. Then put the user in the DSNs: SIGNOZ_TELEMETRYSTORE_CLICKHOUSE_DSN=tcp://signoz:<password>@clickhouse:9000.

Encrypt ClickHouse data at rest

# ClickHouse storage configuration (YAML form)
storage_configuration:
  disks:
    encrypted:
      type: encrypted
      disk: default
      path: encrypted/
      key_hex: "REPLACE_WITH_32_BYTE_HEX_KEY"
  policies:
    encrypted_policy:
      volumes:
        main:
          disk: encrypted

Volume-level encryption (for example, encrypted EBS or LUKS) is a simpler alternative.

Commands & Recipes

Health checks

# Foundry / Docker
docker ps --format 'table {{.Names}}\t{{.Status}}'
curl -fsS http://localhost:8080/api/v1/health
docker exec signoz-telemetrystore-clickhouse-0-0 clickhouse-client --query "SELECT version()"

# Kubernetes (release "signoz" in namespace "signoz")
kubectl -n signoz get pods
kubectl -n signoz exec statefulset/signoz -- wget -qO- http://localhost:8080/api/v1/health
CH_POD=$(kubectl -n signoz get pods -o name | grep '^pod/chi-' | head -1)
kubectl -n signoz exec -it "$CH_POD" -- clickhouse-client --query "SELECT 1"

Logs and restarts

# Docker (Foundry)
docker logs -f signoz-signoz-0
docker compose -f pours/deployment/compose.yaml logs -f ingester

# Kubernetes
kubectl -n signoz logs statefulset/signoz --tail=100
kubectl -n signoz logs deploy/signoz-otel-collector --tail=100
kubectl -n signoz rollout restart statefulset/signoz
kubectl -n signoz rollout restart deploy/signoz-otel-collector

Collector throughput

kubectl -n signoz port-forward deploy/signoz-otel-collector 8888:8888
curl -s http://localhost:8888/metrics | grep -E 'otelcol_receiver_accepted|otelcol_exporter_send_failed'

ClickHouse disk usage by table

SELECT database, table,
       formatReadableSize(sum(bytes_on_disk)) AS size,
       sum(rows) AS rows
FROM system.parts
WHERE active AND database LIKE 'signoz_%'
GROUP BY database, table
ORDER BY sum(bytes_on_disk) DESC
LIMIT 20;

Useful queries (current v3 / v2 schemas)

-- Top services by span count, last hour
SELECT serviceName, count() AS spans
FROM signoz_traces.distributed_signoz_index_v3
WHERE ts_bucket_start >= toUnixTimestamp(now() - INTERVAL 1 HOUR) - 1800
  AND timestamp >= now() - INTERVAL 1 HOUR
GROUP BY serviceName
ORDER BY spans DESC
LIMIT 20;

-- Error ratio by service, last hour
SELECT serviceName,
       countIf(has_error) / count() AS error_ratio
FROM signoz_traces.distributed_signoz_index_v3
WHERE ts_bucket_start >= toUnixTimestamp(now() - INTERVAL 1 HOUR) - 1800
  AND timestamp >= now() - INTERVAL 1 HOUR
GROUP BY serviceName
HAVING count() > 100
ORDER BY error_ratio DESC;

-- Log volume by severity, last hour (timestamp is UInt64 nanoseconds)
SELECT severity_text, count() AS logs
FROM signoz_logs.distributed_logs_v2
WHERE ts_bucket_start >= toUnixTimestamp(now() - INTERVAL 1 HOUR) - 1800
  AND timestamp >= (toUnixTimestamp(now()) - 3600) * 1000000000
GROUP BY severity_text
ORDER BY logs DESC;

-- LLM cost by model, last day (v0.143.0+ attributes)
SELECT attributes_string['gen_ai.request.model'] AS model,
       sum(attributes_number['signoz.gen_ai.usage.tokens.cost']) AS cost
FROM signoz_traces.distributed_signoz_index_v3
WHERE ts_bucket_start >= toUnixTimestamp(now() - INTERVAL 1 DAY) - 1800
  AND timestamp >= now() - INTERVAL 1 DAY
  AND mapContains(attributes_string, 'gen_ai.request.model')
GROUP BY model
ORDER BY cost DESC;

Tip

Add a resource-fingerprint CTE on distributed_traces_v3_resource / distributed_logs_v2_resource whenever you filter by service.name or other resource attributes. The query patterns are in the traces query docs.

Why attributes_number

The collector's signozllmpricingprocessor writes every computed cost attribute with PutDouble, so the cost lands in the numeric attribute map (processor.go, checked 2026-09-27). The output key names come from the pricing-rule config that SigNoz pushes, so confirm the key in the Explorer if you changed the rules.

Troubleshooting

Symptom Likely cause Fix
UI unreachable on 3301 Old port. The UI moved to 8080 in the single-binary architecture. Use http://<host>:8080
SigNoz exits with jwt::secret must be set when provider is jwt JWT provider without a secret (v0.143.0+) Set SIGNOZ_TOKENIZER_JWT_SECRET or drop SIGNOZ_TOKENIZER_PROVIDER=jwt
Collector log unknown type: signozspanmapper SigNoz v0.143.0 with a collector older than v0.144.11 Upgrade the collector together with SigNoz
Collector waits at migrate sync check Migrations not finished or failed Check the signoz-telemetrystore-migrator logs. Re-run cast or the Helm upgrade.
ClickHouse Keeper restart loop (exit 139) on Windows Docker Desktop virtualization Run Docker Engine inside WSL 2
No data after docker compose down -v -v deleted the volumes Restore from backup. Never use -v when migrating.
Queries slow on large ranges No resource filter, or too many parts Add resource filters. Check merges and system.parts.

Sources