Skip to content

Explanation

What this page covers

How SigNoz works inside and why it is built that way. It covers the single SigNoz binary, the OpenTelemetry-based ingestion layer, the ClickHouse storage design, the mutation-aware schema migrator, query execution, alerting, AI observability, the security model and the main design trade-offs. Look-up tables (ports, table names, config keys, pricing, sizing) are in Reference. Step-by-step tasks are in How-to Guides.

Architecture Overview

SigNoz has three planes. The ingestion plane is SigNoz's own distribution of the OpenTelemetry Collector, called the ingester in Foundry. The storage plane is ClickHouse with ClickHouse Keeper or ZooKeeper for coordination, plus a small relational metastore (SQLite or PostgreSQL) for application state. The serving plane is one Go binary, signoz, that bundles the React UI, the API server, the ruler, Alertmanager and an OpAMP server (architecture docs).

The diagram below shows the default self-hosted topology generated by Foundry (v0.143.0), with real service and table names:

flowchart TB
    subgraph Sources["Telemetry sources"]
        SDK["Apps with OTel SDKs<br/>(incl. gen_ai spans)"]
        AGENT["Upstream OTel Collectors<br/>(k8s-infra, host agents)"]
        LEGACY["Jaeger / Zipkin / Prometheus /<br/>FluentBit (optional receivers)"]
    end

    subgraph Ingest["Ingestion plane"]
        ING["ingester<br/>signoz/signoz-otel-collector<br/>OTLP :4317 / :4318"]
        MIG["telemetrystore-migrator<br/>(migrate bootstrap, sync, async)"]
    end

    subgraph Serve["Serving plane: signoz binary :8080"]
        API["API server + querier"]
        UI["React UI (bundled)"]
        RULER["Ruler"]
        AM["Alertmanager"]
        OPAMP["OpAMP server"]
    end

    subgraph Store["Storage plane"]
        CH[("ClickHouse<br/>signoz_traces / signoz_logs /<br/>signoz_metrics / signoz_meter /<br/>signoz_metadata")]
        KEEP["ClickHouse Keeper<br/>or ZooKeeper"]
        META[("Metastore<br/>SQLite or PostgreSQL")]
    end

    SDK -->|OTLP| ING
    AGENT -->|OTLP| ING
    LEGACY -.-> ING
    ING -->|"native TCP :9000"| CH
    MIG -->|DDL| CH
    CH --- KEEP
    API -->|SQL| CH
    RULER -->|"rule queries"| CH
    RULER --> AM
    API --> META
    UI --> API
    OPAMP -.->|"pipelines, span mapper,<br/>LLM pricing config"| ING

Component Responsibility Matrix

Component Language Role Scales via
Ingester (SigNoz OTel Collector) Go Receives OTLP, runs processors, writes to ClickHouse Horizontal replicas behind a load balancer
Agent collectors (optional) Go Per-node collection (k8s-infra chart, host agents) that forward OTLP to the ingester DaemonSet
SigNoz binary Go + TypeScript UI API, querier, UI, ruler, Alertmanager, OpAMP Replicas via signoz.replicaCount. The sharder is a no-op by default, so running several replicas needs care.
Telemetry store migrator Go (collector binary) Schema bootstrap and migrations One-shot job per upgrade
ClickHouse C++ Columnar storage for all signals Shards (write throughput) and replicas (HA, reads)
ClickHouse Keeper / ZooKeeper C++ / Java Replication and ON CLUSTER DDL coordination 3-node ensemble in production
Metastore SQLite / PostgreSQL Users, orgs, dashboards, alert rules, pipelines, pricing rules Postgres HA (managed DB recommended)

Architecture change: from query-service + frontend to one binary

Older SigNoz releases shipped separate query-service, frontend (UI on port 3301) and alertmanager containers. Current releases ship a single signoz/signoz image that serves the UI and API together on port 8080 (image entrypoint ./signoz server). The architecture page describing the bundled binary is dated 2025-04-03. The merge shipped in v0.76.0 (upgrade to v0.76 guide). The pkg/query-service packages still exist in the repository as internal code, but the binary is now built from cmd/community and cmd/enterprise. Commands and dashboards that target signoz-query-service or signoz-frontend are stale.

Ingestion Pipeline

SigNoz OTel Collector distribution

The ingester is a normal OpenTelemetry Collector build: upstream Collector and Collector-Contrib components plus SigNoz-specific receivers, processors, exporters, a connector and a health-check extension (collector README). It runs in one of two modes:

  • Static: --config <file>.
  • Managed: --manager-config <opamp.yaml>. The collector connects to the OpAMP server inside the SigNoz binary, receives its effective config, and writes the active config to --copy-path.

Foundry starts the ingester with both flags. migrate sync check runs first, so an ingester never writes against a schema that has not been migrated yet.

The default traces pipeline (collector v0.144.11) shows where SigNoz adds its own logic:

flowchart LR
    OTLP["otlp receiver"] --> SM["signozspanmetrics/delta<br/>(RED metrics per service/op)"]
    SM --> MAP["signozspanmapper<br/>(normalise to gen_ai.* keys)"]
    MAP --> PRICE["signozllmpricing<br/>(cost per LLM span)"]
    PRICE --> BATCH["batch<br/>50k spans / 5s"]
    BATCH --> CHT["clickhousetraces<br/>signoz_traces"]
    BATCH --> METER["signozmeter connector"]
    BATCH --> MDX["metadataexporter<br/>signoz_metadata"]
    SM -.->|"delta metrics"| CHM["signozclickhousemetrics<br/>signoz_metrics"]
    METER --> CHMETER["signozclickhousemeter<br/>signoz_meter"]
  • signozspanmetrics/delta turns spans into delta-temporality latency histograms and call counts. Dimensions include service.namespace, deployment.environment and service.version, flushed every 60 s. These metrics feed the APM service pages without a separate metrics pipeline.
  • signozspanmapper and signozllmpricing were added in v0.143.0. The span mapper moves framework-specific attributes (for example input.value, gen_ai.prompt) into standard gen_ai.input.messages / gen_ai.output.messages. The pricing processor multiplies gen_ai.usage.*_tokens by per-model prices and writes signoz.gen_ai.usage.*.cost attributes. SigNoz fills both processors' rules over OpAMP.
  • signozmeter counts ingested volume per service, environment and host. It flushes hourly into signoz_meter, which has a 1-year TTL and powers usage and cost-meter views.
  • metadataexporter records attribute keys and value types in signoz_metadata, which drives query-builder autocomplete.

OpAMP-managed configuration

SigNoz uses OpAMP (Open Agent Management Protocol) to push parts of the collector config that users edit in the UI. Today that means log pipelines (parsing and enrichment through signozlogspipelineprocessor), span-mapper groups and LLM pricing rules. The collector applies the new config without the user redeploying it.

What happens when a pushed config is bad: the OpAMP protocol lets the agent report a failed remote-config status and keep running its last effective config. The server sees this in status reports. In SigNoz's server code, a FAILED status marks that config version DeployFailed with the agent's error message, and a success marks it Deployed; the server does not roll back the stored config automatically, but you can redeploy an earlier version (opamp/model/agent.go, agentConf/manager.go, checked 2026-09-27). With several collectors, the code notes that the first status report is the one shown in the UI. Test pipeline changes with the UI's preview before you save them.

Storage Design (ClickHouse)

Why ClickHouse

All three signals, plus meter and metadata, live in one columnar engine. That gives SigNoz:

  • One query language and planner for logs, traces and metrics. Cross-signal joins and correlation happen in SQL instead of across three backends.
  • High-cardinality tolerance. Attributes are stored in typed Map columns and hot keys can be promoted to materialized columns, so there is no per-label inverted index to explode. Unbounded labels on metrics still create time series and remain costly.
  • Compression through per-column codecs: DoubleDelta for timestamps, ZSTD for strings, Gorilla for metric values, T64 for integers.

The trade-off is operational. You run ClickHouse, its Keeper/ZooKeeper quorum and schema migrations yourself, and ClickHouse wants fast local disks and plenty of CPU for background merges.

Resource-fingerprint and time-bucket layout

Logs (logs_v2) and traces (signoz_index_v3) share the same layout. Resource attributes such as service.name and k8s.namespace.name are hashed into a resource_fingerprint and stored once per time bucket in a small companion table (logs_v2_resource, traces_v3_resource). Each row in the main table carries ts_bucket_start and resource_fingerprint, and those two columns lead the sort key.

erDiagram
    traces_v3_resource ||--o{ signoz_index_v3 : "fingerprint = resource_fingerprint"
    logs_v2_resource ||--o{ logs_v2 : "fingerprint = resource_fingerprint"
    signoz_index_v3 ||--o{ logs_v2 : "trace_id correlation"
    traces_v3_resource {
        string labels
        string fingerprint
        int64 seen_at_ts_bucket_start
    }
    signoz_index_v3 {
        uint64 ts_bucket_start
        string resource_fingerprint
        datetime64 timestamp
        fixedstring trace_id
        string span_id
        uint64 duration_nano
        bool has_error
    }
    logs_v2_resource {
        string labels
        string fingerprint
        int64 seen_at_ts_bucket_start
    }
    logs_v2 {
        uint64 ts_bucket_start
        string resource_fingerprint
        uint64 timestamp
        string trace_id
        string body
    }

A query first resolves matching fingerprints from the resource table, then scans only those granules of the main table (resource_fingerprint GLOBAL IN (...) plus a ts_bucket_start range padded by 1,800 s). This is why SigNoz's own docs tell you to always filter on resource attributes (traces query docs). Tables use ttl_only_drop_parts = 1, daily partitions and TTLs in seconds (15 days for logs and traces, 30 days for metrics), so retention drops whole parts instead of running row-level deletes.

Metrics storage

The metrics exporter writes a few landing tables: samples_v4 (values), time_series_v4 (label sets keyed by fingerprint), exp_hist and metadata. Everything else in signoz_metrics is derived by materialized views: time_series_v4_6hrs, _1day and _1week rollups, plus optional reduced tables when metric cardinality reduction rules are enabled (exporter README). Sort keys start with (env, temporality, metric_name, fingerprint, unix_milli). PromQL is supported through a Prometheus-compatible query engine over these tables. Query-builder metrics queries compile to ClickHouse SQL directly.

Schema Migrations

SigNoz first used golang-migrate for ClickHouse schemas and hit three problems: no cluster mode, ClickHouse mutations that time out and leave the database "dirty", and ON CLUSTER DDL piling up in system.distributed_ddl_queue behind a pending mutation. Its replacement, the schema migrator inside signoz-otel-collector, classifies every operation as mutation or not, idempotent or not, and lightweight or not (migrator README).

stateDiagram-v2
    [*] --> Ready: migrate ready
    Ready --> Bootstrap: ClickHouse reachable
    Bootstrap --> SyncUp: migrate bootstrap<br/>(databases, base tables)
    SyncUp --> Serving: migrate sync up<br/>(fast, non-mutating DDL)
    Serving --> AsyncUp: migrate async up<br/>(mutations in background)
    AsyncUp --> Done
    Serving --> Done
    Done --> [*]
    note right of Serving
        ingester runs "migrate sync check"
        before it starts writing
    end note
  • Sync migrations are cheap metadata changes. They must finish before new collectors start.
  • Async migrations (mutations, materializations) run in the background and do not block the upgrade.
  • Replication: --clickhouse-replication / telemetryStoreMigrator.enableReplication makes the migrator emit Replicated* engines and ON CLUSTER DDL.

This is why upgrades have required stops. A release can depend on an async migration from an earlier release having finished.

Query Execution

Signal Query options How it runs
Traces Query builder, ClickHouse SQL Builder compiles to SQL on distributed_signoz_index_v3 with the resource-fingerprint CTE
Logs Query builder, ClickHouse SQL Same pattern on distributed_logs_v2
Metrics Query builder, PromQL, ClickHouse SQL Builder compiles to SQL with temporality-aware functions. PromQL goes through the Prometheus engine.
AI spans AI explorer (v0.141+), query builder Uses normalized gen_ai.* attributes

Query flow for a dashboard panel:

sequenceDiagram
    participant B as Browser (React UI)
    participant S as signoz :8080 (querier)
    participant M as Metastore
    participant C as ClickHouse
    B->>S: POST query range (builder JSON, PromQL or SQL)
    S->>M: Load dashboard, variables, user permissions
    S->>S: Compile builder spec to SQL (resource CTE, ts_bucket pruning)
    S->>C: Resource filter on traces_v3_resource or logs_v2_resource
    C-->>S: Matching fingerprints
    S->>C: Aggregation on distributed main table
    C-->>S: Rows
    S->>S: Post-process (formulas, fill gaps, cache with flux_interval)
    S-->>B: Series or table result

The querier caches results (querier.cache_ttl, default 168h) and re-queries only the most recent flux_interval window (default 5m), because recent data may still be arriving.

Cross-Signal Correlation

Correlation relies on shared identifiers rather than a join service:

flowchart LR
    T["Span<br/>trace_id, span_id,<br/>service.name"] <-->|"trace_id in log"| L["Log<br/>trace_id, span_id"]
    T <-->|"service + time window"| M["Metric<br/>service.name, operation"]
    L <-->|"service + time window"| M
    T -->|"signozspanmetrics"| M
  • Trace to logs: the span view queries logs_v2 by trace_id. The querier pads the time window (log_trace_id_window_padding, default 5m).
  • Logs to trace: a log with trace_id links to the trace waterfall.
  • Metrics to traces: APM charts built from span metrics drill down into the spans for the same service, operation and time range.
  • Exceptions: span events with exception attributes appear in a dedicated Exceptions view.

Alerting Pipeline

The ruler and Alertmanager run inside the SigNoz binary (alertmanager.provider: signoz), so there is no separate Alertmanager deployment to run.

flowchart LR
    R["Alert rule<br/>(builder, PromQL, CH SQL,<br/>anomaly)"] --> E["Ruler<br/>(periodic evaluation)"]
    E -->|"state change"| H[("signoz_analytics<br/>rule_state_history_v0")]
    E -->|"firing / resolved"| A["Alertmanager<br/>(dedupe, group, route)"]
    A --> SL["Slack"]
    A --> PD["PagerDuty"]
    A --> OG["Opsgenie"]
    A --> MT["MS Teams"]
    A --> JI["Jira / incident.io"]
    A --> WH["Webhook / Email"]
  • Rules can target metrics, logs, traces or exceptions.
  • Anomaly-based alerts (z-score with seasonality) are available on Cloud and Enterprise Self-Hosted (docs).
  • Alert state history is stored in ClickHouse, not in the metastore.

AI and LLM Observability

SigNoz treats LLM calls as ordinary OpenTelemetry spans that follow the gen_ai.* semantic conventions, not as a separate product with its own agent (LLM observability overview). v0.143.0 (2026-09-23) added an AI Observability section:

  • Overview: tokens, cost, latency and errors across models, providers and agents.
  • Explorer: queries AI spans by model, provider, agent and tool.
  • Attribute mapping: runs in signozspanmapper and normalizes framework-specific keys into gen_ai.*. By default it moves the key rather than copying it, to avoid storing large prompt payloads twice.
  • Pricing: runs in signozllmpricing and computes cost for every span from per-model prices kept in the metastore.

Server-side inference engines (vLLM, SGLang) expose Prometheus metrics instead of gen_ai spans. The collector scrapes those, so you can place GPU-side signals next to client spans. SigNoz also ships an MCP server (hosted on Cloud, or self-deployed with Foundry on port 8000) for coding agents. Noz is an in-product AI assistant that is available only on SigNoz Cloud.

Security Model

Trust boundaries

flowchart TD
    subgraph External["External"]
        Apps["Instrumented apps"]
        Users["Browser and API clients"]
        IdP["Identity provider<br/>SAML / OIDC / Google"]
    end
    subgraph Ingestion["Ingestion boundary"]
        ING["ingester :4317/:4318<br/>(Cloud: signoz-ingestion-key)"]
    end
    subgraph Platform["SigNoz binary :8080"]
        AUTHN["identn: opaque session tokens,<br/>SIGNOZ-API-KEY"]
        AUTHZ["authz: OpenFGA roles"]
        QRY["API and querier"]
    end
    subgraph Data["Data stores (internal only)"]
        CH[("ClickHouse")]
        META[("Metastore")]
    end
    Apps -->|OTLP| ING
    ING --> CH
    Users -->|"login or API key"| AUTHN
    AUTHN -->|"SSO redirect"| IdP
    AUTHN --> AUTHZ --> QRY
    QRY --> CH
    QRY --> META
Surface Self-hosted Community SigNoz Cloud
OTLP ingestion No authentication by default. Protect it at the network layer. Write-only ingestion keys (signoz-ingestion-key header) with optional per-signal limits
UI sessions Opaque tokens by default since v0.143.0 (JWT optional, secret required) Same
API automation Service accounts, SIGNOZ-API-KEY header Same
SSO Google Workspace OAuth2. SAML/OIDC need Enterprise. SAML, OIDC, Google
Authorization Managed roles (Admin, Editor, Viewer, Anonymous) Plus custom roles and fine-grained transactions (beta, licensed)

Threat-model notes:

  • Open OTLP ports are the main self-hosted risk. Anyone who can reach 4317/4318 can write data, inflate storage or spoof services. Ingestion keys exist only on Cloud (ingestion keys docs).
  • Session tampering. v0.143.0 made SigNoz refuse to start with the JWT provider and no secret, because, in the words of the startup error, "without a JWT secret, user sessions are vulnerable to tampering and unauthorized access".
  • ClickHouse is trusted. SigNoz connects with one DSN, and anyone with ClickHouse access can read all telemetry. Keep 9000/8123 internal and use a dedicated user.
  • Deprovisioning. Removing a user from the IdP blocks SSO login but does not delete the SigNoz account (SSO overview).
  • Prompt data. LLM spans can carry prompts and completions (gen_ai.input.messages). Treat them as sensitive and consider the collector's redaction or transform processors.

The hardening checklist and port table are in Reference. The setup steps are in How-to Guides.

Deployment Topologies

Single node (Foundry Docker Compose)

Foundry's default casting runs one container per molding. This topology fits evaluation and small production workloads where a single ClickHouse node is acceptable:

flowchart LR
    ING["ingester"] --> CH["signoz-telemetrystore-<br/>clickhouse-0-0"]
    MIG["signoz-telemetrystore-migrator"] --> CH
    CH --- KEEP["signoz-telemetrykeeper-<br/>clickhousekeeper-0"]
    SZ["signoz-signoz-0 :8080"] --> CH
    SZ --> PG["signoz-metastore-postgres-0"]
    SZ -.->|OpAMP| ING

Clustered (Kubernetes, sharded ClickHouse)

For higher volume, the Helm chart (or Foundry mode: kubernetes) deploys ClickHouse through the Altinity operator with shards and replicas, a 3-node Keeper/ZooKeeper quorum, a horizontally scaled ingester deployment and optionally PostgreSQL:

flowchart LR
    subgraph Agents["k8s-infra agents (DaemonSet)"]
        A1["otel-agent node 1"]
        A2["otel-agent node N"]
    end
    subgraph Ingesters["signoz-otel-collector Deployment"]
        C1["replica 1"]
        C2["replica 2"]
        C3["replica 3"]
    end
    subgraph CHC["ClickHouse (2 shards x 2 replicas)"]
        S1R1["shard1 rep1"]
        S1R2["shard1 rep2"]
        S2R1["shard2 rep1"]
        S2R2["shard2 rep2"]
    end
    ZK["ZooKeeper x3"]
    SZ["signoz StatefulSet"]
    A1 --> Ingesters
    A2 --> Ingesters
    Ingesters -->|"distributed_* tables"| CHC
    CHC --- ZK
    SZ --> CHC

Writes go to distributed_* tables, which fan out to shards. Replication inside a shard goes through Keeper/ZooKeeper, so coordinator latency directly affects replication lag and ON CLUSTER migrations.

Design Trade-offs

Decision Benefit Cost
One storage engine (ClickHouse) for all signals Single query language, cheap correlation, strong compression You must run ClickHouse well: merges, disks, Keeper, migrations
OTel-native ingestion, no proprietary agent Instrument once, portable data, huge receiver catalog Collector config complexity. Some features need SigNoz-specific processors.
Single signoz binary Fewer moving parts, one port Less independent scaling of UI, API and ruler
Weekly pre-1.0 releases Fast feature delivery (AI observability, Foundry, RBAC) Frequent breaking changes and required upgrade stops
Open core (MIT + ee/) with AGPL collector Free self-hosting of core features SAML/OIDC, custom roles and anomaly alerts are paid. AGPL affects anyone redistributing a modified collector.
Foundry replaces raw Compose files (v0.130.0) One declarative casting for Docker, Swarm, systemd, Kubernetes, ECS and PaaS A new tool to learn. Existing Compose users must migrate.

Benchmarks

The main public performance data is SigNoz's own logs benchmark against ELK and Loki, published in January 2023 (blog, repo). SigNoz ingested about 2.5x faster than ELK with about 50% fewer resources, ran aggregate queries up to 13x faster (ELK won some simple queries), and used about half the storage. Loki could not ingest high-cardinality labels in that setup. The numbers are in Reference.

Read the benchmark critically

  • It is vendor-run and used 2022-era versions (ClickHouse 22.4.5, Elasticsearch 8.4.3, Loki 2.6.1). SigNoz's schema has since moved to logs_v2, and all three systems have changed a lot since.
  • An earlier version of this page said SigNoz compresses "10 to 30x" versus Lucene's "1.5x" and handles "10+ TB/day". No SigNoz source for those figures was found. They have been removed. The benchmark itself reports about 2x less storage than ELK.
  • Real results depend on attribute cardinality, query shape, retention and disk type. Test with your own data.

Known performance considerations

  1. System table growth: ClickHouse system.query_log, part_log and zookeeper_log can grow quickly. Set TTLs on them.
  2. Merges: under heavy ingest, give ClickHouse enough CPU for background merges, or "too many parts" errors will throttle inserts.
  3. Batching: always batch before ClickHouse exporters. The defaults (50k spans / 5 s) follow ClickHouse's advice of large, infrequent inserts.
  4. Resource filters: queries without resource attribute filters scan far more granules.
  5. Metrics cardinality: optional reduction rules (signozclickhousemetrics.reduction) can aggregate away unwanted labels at write time.

Sources