Explanation¶
What this page covers
How SigNoz works inside and why it is built that way. It covers the single SigNoz binary, the OpenTelemetry-based ingestion layer, the ClickHouse storage design, the mutation-aware schema migrator, query execution, alerting, AI observability, the security model and the main design trade-offs. Look-up tables (ports, table names, config keys, pricing, sizing) are in Reference. Step-by-step tasks are in How-to Guides.
Architecture Overview¶
SigNoz has three planes. The ingestion plane is SigNoz's own distribution of the OpenTelemetry Collector, called the ingester in Foundry. The storage plane is ClickHouse with ClickHouse Keeper or ZooKeeper for coordination, plus a small relational metastore (SQLite or PostgreSQL) for application state. The serving plane is one Go binary, signoz, that bundles the React UI, the API server, the ruler, Alertmanager and an OpAMP server (architecture docs).
The diagram below shows the default self-hosted topology generated by Foundry (v0.143.0), with real service and table names:
flowchart TB
subgraph Sources["Telemetry sources"]
SDK["Apps with OTel SDKs<br/>(incl. gen_ai spans)"]
AGENT["Upstream OTel Collectors<br/>(k8s-infra, host agents)"]
LEGACY["Jaeger / Zipkin / Prometheus /<br/>FluentBit (optional receivers)"]
end
subgraph Ingest["Ingestion plane"]
ING["ingester<br/>signoz/signoz-otel-collector<br/>OTLP :4317 / :4318"]
MIG["telemetrystore-migrator<br/>(migrate bootstrap, sync, async)"]
end
subgraph Serve["Serving plane: signoz binary :8080"]
API["API server + querier"]
UI["React UI (bundled)"]
RULER["Ruler"]
AM["Alertmanager"]
OPAMP["OpAMP server"]
end
subgraph Store["Storage plane"]
CH[("ClickHouse<br/>signoz_traces / signoz_logs /<br/>signoz_metrics / signoz_meter /<br/>signoz_metadata")]
KEEP["ClickHouse Keeper<br/>or ZooKeeper"]
META[("Metastore<br/>SQLite or PostgreSQL")]
end
SDK -->|OTLP| ING
AGENT -->|OTLP| ING
LEGACY -.-> ING
ING -->|"native TCP :9000"| CH
MIG -->|DDL| CH
CH --- KEEP
API -->|SQL| CH
RULER -->|"rule queries"| CH
RULER --> AM
API --> META
UI --> API
OPAMP -.->|"pipelines, span mapper,<br/>LLM pricing config"| ING
Component Responsibility Matrix¶
| Component | Language | Role | Scales via |
|---|---|---|---|
| Ingester (SigNoz OTel Collector) | Go | Receives OTLP, runs processors, writes to ClickHouse | Horizontal replicas behind a load balancer |
| Agent collectors (optional) | Go | Per-node collection (k8s-infra chart, host agents) that forward OTLP to the ingester | DaemonSet |
| SigNoz binary | Go + TypeScript UI | API, querier, UI, ruler, Alertmanager, OpAMP | Replicas via signoz.replicaCount. The sharder is a no-op by default, so running several replicas needs care. |
| Telemetry store migrator | Go (collector binary) | Schema bootstrap and migrations | One-shot job per upgrade |
| ClickHouse | C++ | Columnar storage for all signals | Shards (write throughput) and replicas (HA, reads) |
| ClickHouse Keeper / ZooKeeper | C++ / Java | Replication and ON CLUSTER DDL coordination |
3-node ensemble in production |
| Metastore | SQLite / PostgreSQL | Users, orgs, dashboards, alert rules, pipelines, pricing rules | Postgres HA (managed DB recommended) |
Architecture change: from query-service + frontend to one binary
Older SigNoz releases shipped separate query-service, frontend (UI on port 3301) and alertmanager containers. Current releases ship a single signoz/signoz image that serves the UI and API together on port 8080 (image entrypoint ./signoz server). The architecture page describing the bundled binary is dated 2025-04-03. The merge shipped in v0.76.0 (upgrade to v0.76 guide). The pkg/query-service packages still exist in the repository as internal code, but the binary is now built from cmd/community and cmd/enterprise. Commands and dashboards that target signoz-query-service or signoz-frontend are stale.
Ingestion Pipeline¶
SigNoz OTel Collector distribution¶
The ingester is a normal OpenTelemetry Collector build: upstream Collector and Collector-Contrib components plus SigNoz-specific receivers, processors, exporters, a connector and a health-check extension (collector README). It runs in one of two modes:
- Static:
--config <file>. - Managed:
--manager-config <opamp.yaml>. The collector connects to the OpAMP server inside the SigNoz binary, receives its effective config, and writes the active config to--copy-path.
Foundry starts the ingester with both flags. migrate sync check runs first, so an ingester never writes against a schema that has not been migrated yet.
The default traces pipeline (collector v0.144.11) shows where SigNoz adds its own logic:
flowchart LR
OTLP["otlp receiver"] --> SM["signozspanmetrics/delta<br/>(RED metrics per service/op)"]
SM --> MAP["signozspanmapper<br/>(normalise to gen_ai.* keys)"]
MAP --> PRICE["signozllmpricing<br/>(cost per LLM span)"]
PRICE --> BATCH["batch<br/>50k spans / 5s"]
BATCH --> CHT["clickhousetraces<br/>signoz_traces"]
BATCH --> METER["signozmeter connector"]
BATCH --> MDX["metadataexporter<br/>signoz_metadata"]
SM -.->|"delta metrics"| CHM["signozclickhousemetrics<br/>signoz_metrics"]
METER --> CHMETER["signozclickhousemeter<br/>signoz_meter"]
signozspanmetrics/deltaturns spans into delta-temporality latency histograms and call counts. Dimensions includeservice.namespace,deployment.environmentandservice.version, flushed every 60 s. These metrics feed the APM service pages without a separate metrics pipeline.signozspanmapperandsignozllmpricingwere added in v0.143.0. The span mapper moves framework-specific attributes (for exampleinput.value,gen_ai.prompt) into standardgen_ai.input.messages/gen_ai.output.messages. The pricing processor multipliesgen_ai.usage.*_tokensby per-model prices and writessignoz.gen_ai.usage.*.costattributes. SigNoz fills both processors' rules over OpAMP.signozmetercounts ingested volume per service, environment and host. It flushes hourly intosignoz_meter, which has a 1-year TTL and powers usage and cost-meter views.metadataexporterrecords attribute keys and value types insignoz_metadata, which drives query-builder autocomplete.
OpAMP-managed configuration¶
SigNoz uses OpAMP (Open Agent Management Protocol) to push parts of the collector config that users edit in the UI. Today that means log pipelines (parsing and enrichment through signozlogspipelineprocessor), span-mapper groups and LLM pricing rules. The collector applies the new config without the user redeploying it.
What happens when a pushed config is bad: the OpAMP protocol lets the agent report a failed remote-config status and keep running its last effective config. The server sees this in status reports. In SigNoz's server code, a FAILED status marks that config version DeployFailed with the agent's error message, and a success marks it Deployed; the server does not roll back the stored config automatically, but you can redeploy an earlier version (opamp/model/agent.go, agentConf/manager.go, checked 2026-09-27). With several collectors, the code notes that the first status report is the one shown in the UI. Test pipeline changes with the UI's preview before you save them.
Storage Design (ClickHouse)¶
Why ClickHouse¶
All three signals, plus meter and metadata, live in one columnar engine. That gives SigNoz:
- One query language and planner for logs, traces and metrics. Cross-signal joins and correlation happen in SQL instead of across three backends.
- High-cardinality tolerance. Attributes are stored in typed
Mapcolumns and hot keys can be promoted to materialized columns, so there is no per-label inverted index to explode. Unbounded labels on metrics still create time series and remain costly. - Compression through per-column codecs:
DoubleDeltafor timestamps,ZSTDfor strings,Gorillafor metric values,T64for integers.
The trade-off is operational. You run ClickHouse, its Keeper/ZooKeeper quorum and schema migrations yourself, and ClickHouse wants fast local disks and plenty of CPU for background merges.
Resource-fingerprint and time-bucket layout¶
Logs (logs_v2) and traces (signoz_index_v3) share the same layout. Resource attributes such as service.name and k8s.namespace.name are hashed into a resource_fingerprint and stored once per time bucket in a small companion table (logs_v2_resource, traces_v3_resource). Each row in the main table carries ts_bucket_start and resource_fingerprint, and those two columns lead the sort key.
erDiagram
traces_v3_resource ||--o{ signoz_index_v3 : "fingerprint = resource_fingerprint"
logs_v2_resource ||--o{ logs_v2 : "fingerprint = resource_fingerprint"
signoz_index_v3 ||--o{ logs_v2 : "trace_id correlation"
traces_v3_resource {
string labels
string fingerprint
int64 seen_at_ts_bucket_start
}
signoz_index_v3 {
uint64 ts_bucket_start
string resource_fingerprint
datetime64 timestamp
fixedstring trace_id
string span_id
uint64 duration_nano
bool has_error
}
logs_v2_resource {
string labels
string fingerprint
int64 seen_at_ts_bucket_start
}
logs_v2 {
uint64 ts_bucket_start
string resource_fingerprint
uint64 timestamp
string trace_id
string body
}
A query first resolves matching fingerprints from the resource table, then scans only those granules of the main table (resource_fingerprint GLOBAL IN (...) plus a ts_bucket_start range padded by 1,800 s). This is why SigNoz's own docs tell you to always filter on resource attributes (traces query docs). Tables use ttl_only_drop_parts = 1, daily partitions and TTLs in seconds (15 days for logs and traces, 30 days for metrics), so retention drops whole parts instead of running row-level deletes.
Metrics storage¶
The metrics exporter writes a few landing tables: samples_v4 (values), time_series_v4 (label sets keyed by fingerprint), exp_hist and metadata. Everything else in signoz_metrics is derived by materialized views: time_series_v4_6hrs, _1day and _1week rollups, plus optional reduced tables when metric cardinality reduction rules are enabled (exporter README). Sort keys start with (env, temporality, metric_name, fingerprint, unix_milli). PromQL is supported through a Prometheus-compatible query engine over these tables. Query-builder metrics queries compile to ClickHouse SQL directly.
Schema Migrations¶
SigNoz first used golang-migrate for ClickHouse schemas and hit three problems: no cluster mode, ClickHouse mutations that time out and leave the database "dirty", and ON CLUSTER DDL piling up in system.distributed_ddl_queue behind a pending mutation. Its replacement, the schema migrator inside signoz-otel-collector, classifies every operation as mutation or not, idempotent or not, and lightweight or not (migrator README).
stateDiagram-v2
[*] --> Ready: migrate ready
Ready --> Bootstrap: ClickHouse reachable
Bootstrap --> SyncUp: migrate bootstrap<br/>(databases, base tables)
SyncUp --> Serving: migrate sync up<br/>(fast, non-mutating DDL)
Serving --> AsyncUp: migrate async up<br/>(mutations in background)
AsyncUp --> Done
Serving --> Done
Done --> [*]
note right of Serving
ingester runs "migrate sync check"
before it starts writing
end note
- Sync migrations are cheap metadata changes. They must finish before new collectors start.
- Async migrations (mutations, materializations) run in the background and do not block the upgrade.
- Replication:
--clickhouse-replication/telemetryStoreMigrator.enableReplicationmakes the migrator emitReplicated*engines andON CLUSTERDDL.
This is why upgrades have required stops. A release can depend on an async migration from an earlier release having finished.
Query Execution¶
| Signal | Query options | How it runs |
|---|---|---|
| Traces | Query builder, ClickHouse SQL | Builder compiles to SQL on distributed_signoz_index_v3 with the resource-fingerprint CTE |
| Logs | Query builder, ClickHouse SQL | Same pattern on distributed_logs_v2 |
| Metrics | Query builder, PromQL, ClickHouse SQL | Builder compiles to SQL with temporality-aware functions. PromQL goes through the Prometheus engine. |
| AI spans | AI explorer (v0.141+), query builder | Uses normalized gen_ai.* attributes |
Query flow for a dashboard panel:
sequenceDiagram
participant B as Browser (React UI)
participant S as signoz :8080 (querier)
participant M as Metastore
participant C as ClickHouse
B->>S: POST query range (builder JSON, PromQL or SQL)
S->>M: Load dashboard, variables, user permissions
S->>S: Compile builder spec to SQL (resource CTE, ts_bucket pruning)
S->>C: Resource filter on traces_v3_resource or logs_v2_resource
C-->>S: Matching fingerprints
S->>C: Aggregation on distributed main table
C-->>S: Rows
S->>S: Post-process (formulas, fill gaps, cache with flux_interval)
S-->>B: Series or table result
The querier caches results (querier.cache_ttl, default 168h) and re-queries only the most recent flux_interval window (default 5m), because recent data may still be arriving.
Cross-Signal Correlation¶
Correlation relies on shared identifiers rather than a join service:
flowchart LR
T["Span<br/>trace_id, span_id,<br/>service.name"] <-->|"trace_id in log"| L["Log<br/>trace_id, span_id"]
T <-->|"service + time window"| M["Metric<br/>service.name, operation"]
L <-->|"service + time window"| M
T -->|"signozspanmetrics"| M
- Trace to logs: the span view queries
logs_v2bytrace_id. The querier pads the time window (log_trace_id_window_padding, default 5m). - Logs to trace: a log with
trace_idlinks to the trace waterfall. - Metrics to traces: APM charts built from span metrics drill down into the spans for the same service, operation and time range.
- Exceptions: span events with exception attributes appear in a dedicated Exceptions view.
Alerting Pipeline¶
The ruler and Alertmanager run inside the SigNoz binary (alertmanager.provider: signoz), so there is no separate Alertmanager deployment to run.
flowchart LR
R["Alert rule<br/>(builder, PromQL, CH SQL,<br/>anomaly)"] --> E["Ruler<br/>(periodic evaluation)"]
E -->|"state change"| H[("signoz_analytics<br/>rule_state_history_v0")]
E -->|"firing / resolved"| A["Alertmanager<br/>(dedupe, group, route)"]
A --> SL["Slack"]
A --> PD["PagerDuty"]
A --> OG["Opsgenie"]
A --> MT["MS Teams"]
A --> JI["Jira / incident.io"]
A --> WH["Webhook / Email"]
- Rules can target metrics, logs, traces or exceptions.
- Anomaly-based alerts (z-score with seasonality) are available on Cloud and Enterprise Self-Hosted (docs).
- Alert state history is stored in ClickHouse, not in the metastore.
AI and LLM Observability¶
SigNoz treats LLM calls as ordinary OpenTelemetry spans that follow the gen_ai.* semantic conventions, not as a separate product with its own agent (LLM observability overview). v0.143.0 (2026-09-23) added an AI Observability section:
- Overview: tokens, cost, latency and errors across models, providers and agents.
- Explorer: queries AI spans by model, provider, agent and tool.
- Attribute mapping: runs in
signozspanmapperand normalizes framework-specific keys intogen_ai.*. By default it moves the key rather than copying it, to avoid storing large prompt payloads twice. - Pricing: runs in
signozllmpricingand computes cost for every span from per-model prices kept in the metastore.
Server-side inference engines (vLLM, SGLang) expose Prometheus metrics instead of gen_ai spans. The collector scrapes those, so you can place GPU-side signals next to client spans. SigNoz also ships an MCP server (hosted on Cloud, or self-deployed with Foundry on port 8000) for coding agents. Noz is an in-product AI assistant that is available only on SigNoz Cloud.
Security Model¶
Trust boundaries¶
flowchart TD
subgraph External["External"]
Apps["Instrumented apps"]
Users["Browser and API clients"]
IdP["Identity provider<br/>SAML / OIDC / Google"]
end
subgraph Ingestion["Ingestion boundary"]
ING["ingester :4317/:4318<br/>(Cloud: signoz-ingestion-key)"]
end
subgraph Platform["SigNoz binary :8080"]
AUTHN["identn: opaque session tokens,<br/>SIGNOZ-API-KEY"]
AUTHZ["authz: OpenFGA roles"]
QRY["API and querier"]
end
subgraph Data["Data stores (internal only)"]
CH[("ClickHouse")]
META[("Metastore")]
end
Apps -->|OTLP| ING
ING --> CH
Users -->|"login or API key"| AUTHN
AUTHN -->|"SSO redirect"| IdP
AUTHN --> AUTHZ --> QRY
QRY --> CH
QRY --> META
| Surface | Self-hosted Community | SigNoz Cloud |
|---|---|---|
| OTLP ingestion | No authentication by default. Protect it at the network layer. | Write-only ingestion keys (signoz-ingestion-key header) with optional per-signal limits |
| UI sessions | Opaque tokens by default since v0.143.0 (JWT optional, secret required) | Same |
| API automation | Service accounts, SIGNOZ-API-KEY header |
Same |
| SSO | Google Workspace OAuth2. SAML/OIDC need Enterprise. | SAML, OIDC, Google |
| Authorization | Managed roles (Admin, Editor, Viewer, Anonymous) | Plus custom roles and fine-grained transactions (beta, licensed) |
Threat-model notes:
- Open OTLP ports are the main self-hosted risk. Anyone who can reach 4317/4318 can write data, inflate storage or spoof services. Ingestion keys exist only on Cloud (ingestion keys docs).
- Session tampering. v0.143.0 made SigNoz refuse to start with the JWT provider and no secret, because, in the words of the startup error, "without a JWT secret, user sessions are vulnerable to tampering and unauthorized access".
- ClickHouse is trusted. SigNoz connects with one DSN, and anyone with ClickHouse access can read all telemetry. Keep 9000/8123 internal and use a dedicated user.
- Deprovisioning. Removing a user from the IdP blocks SSO login but does not delete the SigNoz account (SSO overview).
- Prompt data. LLM spans can carry prompts and completions (
gen_ai.input.messages). Treat them as sensitive and consider the collector'sredactionortransformprocessors.
The hardening checklist and port table are in Reference. The setup steps are in How-to Guides.
Deployment Topologies¶
Single node (Foundry Docker Compose)¶
Foundry's default casting runs one container per molding. This topology fits evaluation and small production workloads where a single ClickHouse node is acceptable:
flowchart LR
ING["ingester"] --> CH["signoz-telemetrystore-<br/>clickhouse-0-0"]
MIG["signoz-telemetrystore-migrator"] --> CH
CH --- KEEP["signoz-telemetrykeeper-<br/>clickhousekeeper-0"]
SZ["signoz-signoz-0 :8080"] --> CH
SZ --> PG["signoz-metastore-postgres-0"]
SZ -.->|OpAMP| ING
Clustered (Kubernetes, sharded ClickHouse)¶
For higher volume, the Helm chart (or Foundry mode: kubernetes) deploys ClickHouse through the Altinity operator with shards and replicas, a 3-node Keeper/ZooKeeper quorum, a horizontally scaled ingester deployment and optionally PostgreSQL:
flowchart LR
subgraph Agents["k8s-infra agents (DaemonSet)"]
A1["otel-agent node 1"]
A2["otel-agent node N"]
end
subgraph Ingesters["signoz-otel-collector Deployment"]
C1["replica 1"]
C2["replica 2"]
C3["replica 3"]
end
subgraph CHC["ClickHouse (2 shards x 2 replicas)"]
S1R1["shard1 rep1"]
S1R2["shard1 rep2"]
S2R1["shard2 rep1"]
S2R2["shard2 rep2"]
end
ZK["ZooKeeper x3"]
SZ["signoz StatefulSet"]
A1 --> Ingesters
A2 --> Ingesters
Ingesters -->|"distributed_* tables"| CHC
CHC --- ZK
SZ --> CHC
Writes go to distributed_* tables, which fan out to shards. Replication inside a shard goes through Keeper/ZooKeeper, so coordinator latency directly affects replication lag and ON CLUSTER migrations.
Design Trade-offs¶
| Decision | Benefit | Cost |
|---|---|---|
| One storage engine (ClickHouse) for all signals | Single query language, cheap correlation, strong compression | You must run ClickHouse well: merges, disks, Keeper, migrations |
| OTel-native ingestion, no proprietary agent | Instrument once, portable data, huge receiver catalog | Collector config complexity. Some features need SigNoz-specific processors. |
Single signoz binary |
Fewer moving parts, one port | Less independent scaling of UI, API and ruler |
| Weekly pre-1.0 releases | Fast feature delivery (AI observability, Foundry, RBAC) | Frequent breaking changes and required upgrade stops |
Open core (MIT + ee/) with AGPL collector |
Free self-hosting of core features | SAML/OIDC, custom roles and anomaly alerts are paid. AGPL affects anyone redistributing a modified collector. |
| Foundry replaces raw Compose files (v0.130.0) | One declarative casting for Docker, Swarm, systemd, Kubernetes, ECS and PaaS | A new tool to learn. Existing Compose users must migrate. |
Benchmarks¶
The main public performance data is SigNoz's own logs benchmark against ELK and Loki, published in January 2023 (blog, repo). SigNoz ingested about 2.5x faster than ELK with about 50% fewer resources, ran aggregate queries up to 13x faster (ELK won some simple queries), and used about half the storage. Loki could not ingest high-cardinality labels in that setup. The numbers are in Reference.
Read the benchmark critically
- It is vendor-run and used 2022-era versions (ClickHouse 22.4.5, Elasticsearch 8.4.3, Loki 2.6.1). SigNoz's schema has since moved to
logs_v2, and all three systems have changed a lot since. - An earlier version of this page said SigNoz compresses "10 to 30x" versus Lucene's "1.5x" and handles "10+ TB/day". No SigNoz source for those figures was found. They have been removed. The benchmark itself reports about 2x less storage than ELK.
- Real results depend on attribute cardinality, query shape, retention and disk type. Test with your own data.
Known performance considerations¶
- System table growth: ClickHouse
system.query_log,part_logandzookeeper_logcan grow quickly. Set TTLs on them. - Merges: under heavy ingest, give ClickHouse enough CPU for background merges, or "too many parts" errors will throttle inserts.
- Batching: always batch before ClickHouse exporters. The defaults (50k spans / 5 s) follow ClickHouse's advice of large, infrequent inserts.
- Resource filters: queries without resource attribute filters scan far more granules.
- Metrics cardinality: optional reduction rules (
signozclickhousemetrics.reduction) can aggregate away unwanted labels at write time.