Explanation¶
What this page covers
How the Victoria Stack works and why it is built that way: topology, the shared insert/select/storage pattern, storage engines (and why there is no WAL), the write and read paths, VictoriaTraces on top of the VictoriaLogs engine, the security and multitenancy model, and the open-core split. For exact flags, ports and versions see Reference; for recipes see How-to guides; hub: Victoria Stack.
Default Topology¶
The reference production topology puts collectors in front of a single vmauth entry point that routes by URL path to three independent storage systems, with Grafana and vmalert as consumers:
flowchart TB
subgraph Sources["Data Sources"]
K8s["Kubernetes<br/>Pods and Services"]
Apps["Applications<br/>(OTel SDK)"]
Infra["Exporters<br/>(node_exporter, etc.)"]
Logs["Log shippers<br/>(Fluent Bit, vlagent, Vector)"]
end
subgraph Collection["Collection Layer"]
Agent["vmagent<br/>scrape + remote write"]
OTel["OTel Collector<br/>(optional)"]
end
subgraph Proxy["Routing Layer"]
Auth["vmauth<br/>auth, route, load balance"]
end
subgraph MetricsCluster["VictoriaMetrics (metrics)"]
MI["vminsert x2"]
MS["vmstorage x3<br/>(StatefulSet, SSD)"]
MSel["vmselect x2"]
MI --> MS
MSel --> MS
end
subgraph LogsCluster["VictoriaLogs (logs)"]
LI["vlinsert x2"]
LS["vlstorage x3"]
LSel["vlselect x2"]
LI --> LS
LSel --> LS
end
subgraph TracesCluster["VictoriaTraces (traces)"]
TI["vtinsert x2"]
TS["vtstorage x3"]
TSel["vtselect x2"]
TI --> TS
TSel --> TS
end
subgraph Alerting["Alerting"]
Alert["vmalert"]
AM["Alertmanager"]
end
subgraph Viz["Visualization"]
Grafana["Grafana"]
VMUI["VMUI<br/>(built-in)"]
end
K8s --> Agent
Infra --> Agent
Apps --> OTel
Agent --> Auth
OTel --> Auth
Logs --> Auth
Auth -->|"/api/v1/write"| MI
Auth -->|"/insert/jsonline, /insert/loki"| LI
Auth -->|"/insert/opentelemetry/v1/traces"| TI
Auth -->|"PromQL / MetricsQL"| MSel
Auth -->|"LogsQL"| LSel
Auth -->|"/select/jaeger, /select/tempo"| TSel
Grafana --> Auth
VMUI --> Auth
Alert --> Auth
Alert -->|"fire alerts"| AM
style Sources fill:#0d7377,color:#fff
style Collection fill:#ff6600,color:#fff
style Proxy fill:#7b42bc,color:#fff
style MetricsCluster fill:#2a2d3e,color:#fff
style LogsCluster fill:#2a7de1,color:#fff
style TracesCluster fill:#e65100,color:#fff
style Alerting fill:#c62828,color:#fff
style Viz fill:#ff6600,color:#fff
The three databases share design ideas and tooling (vmauth, vmalert, VMUI, operator) but not storage: each has its own storage nodes, retention and backups.
Deployment Modes¶
Single-Node vs Cluster¶
| Feature | Single-Node | Cluster |
|---|---|---|
| Scalability | Vertical only | Horizontal and vertical |
| Operational Complexity | Very low (1 binary) | Moderate (3 roles) |
| Multi-tenancy | No (VictoriaMetrics); yes for VictoriaLogs/VictoriaTraces via headers | Yes (accountID:projectID) |
| Replication | No (relies on durable disk) | VictoriaMetrics: -replicationFactor=N; VictoriaLogs/VictoriaTraces: none (shard only) |
| Target Workload | Upstream recommends single-node below ~1M samples/s | Beyond one machine |
| External Dependencies | None | None |
Recommendation (upstream): start with single-node. Move to cluster only for multitenancy, horizontal scale beyond one machine, or application-level replication. A single-node VictoriaLogs or VictoriaTraces instance can later be listed as a storage node of a cluster, so the migration path is incremental.
Component Roles¶
Each signal follows the same tri-role pattern in cluster mode:
| Role | Metrics | Logs | Traces |
|---|---|---|---|
| Ingestion (stateless) | vminsert |
vlinsert |
vtinsert |
| Querying (stateless) | vmselect |
vlselect |
vtselect |
| Storage (stateful) | vmstorage |
vlstorage |
vtstorage |
| Binaries | Three separate binaries | One victoria-logs binary; role set by -storageNode, -insert.disable, -select.disable |
One victoria-traces binary, same model |
| Sharding key | Consistent hash of metric name + labels | Even spread of log streams | Trace ID |
Insert and select roles are stateless and scale on CPU; storage nodes hold the data and scale on disk and memory. vmstorage nodes never talk to each other (shared-nothing), which removes the need for a consensus protocol.
Core Mechanisms¶
Storage Engine (VictoriaMetrics)¶
- Buffer, then immutable parts. Incoming samples are buffered in memory for about a second, turned into searchable in-memory parts, and periodically persisted (
-inmemoryDataFlushInterval) topartdirectories under<storageDataPath>/data/small/YYYY_MM/. Background merges combine small parts into bigger ones (data/big/), an LSM-like design. - No write-ahead log. VictoriaMetrics deliberately has no WAL. On an unclean shutdown (OOM,
kill -9, power loss) the last few seconds of unflushed data can be lost. Upstream's argument: senders (Prometheus, vmagent with its on-disk queue) already retry, and a WAL costs disk IO on every write. See this article. - Blocks by TSID. Each part stores blocks of up to 8K samples sorted by internal
TSID; timestamps and values are encoded separately and compressed (ZSTD plus type-specific encodings). - Per-partition IndexDB. The inverted index (label pairs -> TSID) has global and per-day variants. Since v1.133.0 (2026-01-02) every monthly partition has its own IndexDB, which is dropped with the partition when it leaves retention. The upgrade re-registers all active series once.
- Deterministic sharding. In cluster mode
vminsertuses consistent hashing over metric name and labels to pick vmstorage nodes. Since v1.149.0 slowness-based rerouting is on by default: if one vmstorage node is slow, new series are routed elsewhere (disabled automatically when-replicationFactor> 1).
VictoriaMetrics Write and Query Internals¶
A compact view of how the single-node engine connects ingestion APIs, storage and the query engine:
flowchart LR
subgraph Ingestion["Ingestion APIs"]
PR["Prometheus<br/>remote_write"]
IL["InfluxDB<br/>line protocol"]
DD["Datadog / NewRelic"]
OT["OTLP metrics"]
GR["Graphite"]
end
subgraph VM["VictoriaMetrics storage"]
WB["In-memory buffer<br/>(about 1s)"]
IMP["In-memory parts<br/>(searchable)"]
SMALL["data/small/YYYY_MM<br/>parts"]
BIG["data/big/YYYY_MM<br/>merged parts"]
IDX["Per-partition IndexDB<br/>(labels to TSID)"]
end
subgraph Query["Query"]
MQL["MetricsQL engine"]
end
Ingestion --> WB --> IMP --> SMALL -->|"background merge"| BIG
WB --> IDX
MQL --> IDX
MQL --> IMP
MQL --> SMALL
MQL --> BIG
MQL --> Grafana["Grafana / VMUI"]
style VM fill:#2a2d3e,color:#fff
style Ingestion fill:#0d7377,color:#fff
style Query fill:#ff6600,color:#fff
VictoriaLogs (Logs)¶
VictoriaLogs accepts logs as JSON-like entries with _msg, _time and _stream plus arbitrary fields:
flowchart LR
subgraph Ingestion["Ingestion APIs"]
LK["Loki push"]
ES["Elasticsearch bulk"]
SL["Syslog"]
OT["OTLP logs"]
JL["JSON lines"]
end
subgraph VL["VictoriaLogs storage"]
direction TB
Parse["Parse fields<br/>and _stream"]
Blocks["Per-field column blocks<br/>grouped by stream and time"]
BF["Bloom filters per block<br/>(word and phrase skip)"]
Parts["partitions/YYYYMMDD<br/>(per-day)"]
end
Ingestion --> Parse --> Blocks --> Parts
Blocks --> BF
LogsQL["LogsQL engine<br/>(parallel per CPU core)"] --> BF
LogsQL --> Parts
LogsQL --> Grafana["Grafana plugin / VMUI"]
style VL fill:#2a7de1,color:#fff
style Ingestion fill:#0d7377,color:#fff
- Columnar by field. Values of the same field across entries are stored in one block, so queries read only the fields they need. The design is inspired by ClickHouse.
- Bloom filters, not an inverted index. Word and phrase filters consult per-block bloom filters to skip blocks that cannot match, then scan the rest in parallel on all CPU cores. This keeps RAM low; the trade-off is scan CPU for unselective queries.
- Stream locality. Entries with the same
_streamare stored together, which improves compression and lets stream filters ({app="nginx"}) skip unrelated blocks. A sparse index on_timespeeds time filters. - Per-day partitions at
<storageDataPath>/partitions/YYYYMMDDmake retention a directory delete and allow attach/detach of partitions and per-partition snapshots for backup.
VictoriaTraces (Traces)¶
VictoriaTraces converts OTLP spans into structured log entries and stores them with the VictoriaLogs engine:
flowchart LR
subgraph Ingestion["Ingestion (OTLP only)"]
OTLP_H["OTLP/HTTP<br/>/insert/opentelemetry/v1/traces<br/>:10428"]
OTLP_G["OTLP/gRPC<br/>:4317 when<br/>-otlpGRPCListenAddr set"]
end
subgraph VT["VictoriaTraces"]
direction TB
TP["Span to structured<br/>log entry"]
VLS["VictoriaLogs<br/>storage engine"]
SG["Service graph<br/>relations"]
end
Ingestion --> TP --> VLS
TP --> SG
JQ["Jaeger query API<br/>/select/jaeger"] --> VLS
TQ["Tempo API (experimental)<br/>/select/tempo"] --> VLS
LQ["LogsQL<br/>/select/logsql"] --> VLS
JQ --> Grafana["Grafana<br/>(Jaeger or Tempo data source)"]
TQ --> Grafana
style VT fill:#e65100,color:#fff
style Ingestion fill:#0d7377,color:#fff
- Inherits the VictoriaLogs engine (columnar blocks, bloom filters, per-day partitions, compression) and therefore needs no object storage.
- Ingestion is OTLP only. Jaeger/Zipkin clients should export through the OpenTelemetry Collector. gRPC is off by default and uses TLS unless
-otlpGRPC.tls=false. - Query surfaces. Jaeger query JSON API (Grafana Jaeger data source, Jaeger UI), LogsQL over spans, and an experimental subset of the Grafana Tempo HTTP API under
/select/tempo: search with TraceQL, tag autocomplete, trace by ID (v1 and v2), and/api/metrics/query_rangefor TraceQL metrics. Tempo support started in v0.8.0 (2026-03-02), was extended for Grafana Traces Drilldown in v0.9.0 (2026-05-20), and upstream lists v0.9.4 as the minimum version for it. Some TraceQL functions and drilldown panels are still unsupported. - Maturity. VictoriaTraces is still 0.x (v0.11.1, 2026-09-16). Upstream claims up to 3.7x less RAM and 2.6x less CPU than Tempo, but there is no independent public benchmark at 100M+ spans/day.
Data Flow¶
The cluster write and read paths for metrics (logs and traces follow the same shape with their own roles):
sequenceDiagram
participant App as Targets / Apps
participant Agent as vmagent
participant Auth as vmauth
participant Ins as vminsert
participant Sto as vmstorage nodes
participant Sel as vmselect
participant G as Grafana
Agent->>App: Scrape /metrics
Agent->>Agent: Relabel, buffer to on-disk queue if remote is down
Agent->>Auth: Remote write (Prometheus RW)
Auth->>Auth: Authenticate, pick url_map route
Auth->>Ins: /insert/0/prometheus/api/v1/write
Ins->>Sto: Consistent hash per series, RF copies
Note over Sto: Buffer about 1s, in-memory parts, flush, background merge
G->>Auth: PromQL or MetricsQL query
Auth->>Sel: /select/0/prometheus/api/v1/query_range
Sel->>Sto: Fetch blocks for matching TSIDs from all nodes
Sto-->>Sel: Compressed blocks
Sel->>Sel: Deduplicate, merge, evaluate MetricsQL
Sel-->>G: Result (may be marked partial)
vmagentscrapes targets (or receives pushes), applies relabeling and stream aggregation, and buffers to-remoteWrite.tmpDataPathif the destination is unavailable.vmauthauthenticates the request and routes it by path, host, header or JWT claim to the right backend and tenant.vminsertshards series across vmstorage nodes (N copies with replication).- vmstorage buffers, flushes and merges as described above; there is no WAL replay on restart.
vmselectqueries all vmstorage nodes, deduplicates replicated samples (-dedup.minScrapeInterval), and evaluates MetricsQL. If some nodes are down it returns partial results unless fewer than-replicationFactornodes are missing.
vmalert Evaluation Flow¶
vmalert periodically evaluates rule groups against a datasource (VictoriaMetrics, VictoriaLogs or VictoriaTraces) and writes results back:
sequenceDiagram
participant A as vmalert
participant P as vmauth
participant DS as VictoriaMetrics or VictoriaLogs
participant RW as Remote write target
participant AM as Alertmanager
Note over A: Evaluate rule group every interval
A->>P: Instant query (MetricsQL or LogsQL stats)
P->>DS: Route by path
DS-->>P: Result
P-->>A: Result
alt Alerting rule fires
A->>AM: Send alert notification
A->>RW: Write ALERTS and ALERTS_FOR_STATE series
else Recording rule
A->>RW: Remote write recorded series
end
Alert state is restored after restart from the ALERTS_FOR_STATE series (-remoteRead.url); since tip after v1.152.0 alerts that can be restored straight to firing skip the pending state.
Multi-Source Log Ingestion¶
VictoriaLogs accepts common shipper protocols directly, so no translation tier is required:
flowchart LR
A["Promtail / Alloy"] -->|"Loki push API"| B{"vmauth"}
C["Fluent Bit"] -->|"JSON lines"| B
D["Logstash / Vector"] -->|"Elasticsearch bulk"| B
E["OTel Collector"] -->|"OTLP logs"| B
F["rsyslog"] -->|"Syslog"| G
V["vlagent"] -->|"native, replicated"| B
B -->|"route and set tenant headers"| G["vlinsert"]
G --> H[("vlstorage")]
style B fill:#7b42bc,color:#fff
style H fill:#2a7de1,color:#fff
Syslog is a TCP/UDP listener on VictoriaLogs itself (not HTTP), so it bypasses vmauth.
Storage Layout¶
VictoriaMetrics (Metrics)¶
<storageDataPath>/
├── data/
│ ├── small/YYYY_MM/ # freshly flushed parts, parts.json per partition
│ └── big/YYYY_MM/ # merged large parts
├── indexdb/ # inverted index (per-partition since v1.133.0)
├── snapshots/ # instant snapshots used by vmbackup
└── metadata/, cache/ # internal state and persisted caches
The metadata/ and cache/ entries and the exact on-disk location of per-partition IndexDB after v1.133.0 are not spelled out in the public storage docs; treat that part of the tree as indicative.
VictoriaLogs (Logs)¶
<storageDataPath>/
└── partitions/
└── YYYYMMDD/ # one directory per day; snapshot, attach, detach, delete as a unit
VictoriaTraces (Traces)¶
Same layout as VictoriaLogs (per-day partitions); spans are organized by trace ID and span attributes.
Key Design Decisions¶
| Decision | Rationale | Trade-off |
|---|---|---|
| No external dependencies | No PostgreSQL, Redis, ZooKeeper or object storage — small operational surface | Durability depends on the disks and on backups |
| Local disk over object storage | Lower latency, no request costs | Capacity bound to node disks; DR needs vmbackup/rsync |
| No WAL | Less write IO; senders already retry | Seconds of data can be lost on crash |
| Shared-nothing cluster | Storage nodes do not coordinate — simple scaling | Rebalancing is not automatic; old data stays where it was written |
| Consistent hashing, no consensus | Deterministic placement without Raft/Paxos | Replication is optional and multiplies cost by RF |
| Bloom filters (logs, traces) | Much less RAM than inverted indexes | More CPU for unselective full-text queries |
| Same binary, multiple roles (logs, traces) | A single node can become a storage node of a cluster | Role misconfiguration is possible; disable unused APIs |
| Apache 2.0 + Enterprise features | Permissive core; revenue from downsampling, security and support features | Some production features (downsampling, mTLS, LTS) are paid |
Open-Core Model¶
All three databases and the vm* tools are Apache 2.0. Enterprise builds (same code base plus closed features, -enterprise image tags) add storage-cost features (downsampling, retention filters), operational automation (vmbackupmanager, automatic vmstorage discovery), security features (vmgateway, mTLS, IP filters, auto-TLS, FIPS builds), vmanomaly, and LTS release lines. Since early 2026 several auth features that previously pushed users towards vmgateway landed in community vmauth: JWT verification (v1.137.0), OIDC discovery and claim matching (v1.138.0), and browser SSO (in tip after v1.152.0). The full matrix is in Reference: Community vs Enterprise.
Security Architecture Overview¶
VictoriaMetrics components have no built-in per-user authorization. vminsert, vmselect and vmstorage must run in a private network, and all external access goes through an auth proxy: vmauth (community) or vmgateway (Enterprise).
flowchart TD
subgraph External["External"]
Clients["API clients<br/>Grafana, vmagent"]
Users["Browser users"]
end
subgraph AuthLayer["Auth proxy layer"]
Vmauth["vmauth<br/>Basic, Bearer, JWT/OIDC,<br/>mTLS routing (Enterprise)"]
Vmgateway["vmgateway (Enterprise)<br/>JWT vm_access claims,<br/>rate limits"]
end
subgraph Cluster["VictoriaMetrics cluster (private network)"]
Vminsert["vminsert :8480"]
Vmselect["vmselect :8481"]
Vmstorage["vmstorage :8482"]
end
Clients -->|"Bearer token or Basic auth"| Vmauth
Users -->|"OIDC ID token or SSO cookie"| Vmauth
Users -->|"OIDC JWT"| Vmgateway
Vmauth -->|"route to tenant path or headers"| Vminsert
Vmauth -->|"route to tenant path or headers"| Vmselect
Vmgateway -->|"tenant from vm_access claim"| Vminsert
Vmgateway -->|"tenant from vm_access claim"| Vmselect
Vminsert -->|":8400"| Vmstorage
Vmselect -->|":8401"| Vmstorage
The JWT, OIDC discovery and match_claims capabilities in community vmauth overlap with what used to require vmgateway; vmgateway still adds rate limiting and is Enterprise-only.
Network Topology¶
| Component | Network Zone | Exposure |
|---|---|---|
vmstorage / vlstorage / vtstorage |
Private subnet | No external access |
vminsert / vlinsert / vtinsert |
Private subnet | Behind vmauth only |
vmselect / vlselect / vtselect |
Private subnet | Behind vmauth only |
vmauth |
DMZ / public subnet | TLS termination point |
Multi-Tenant Data Isolation¶
In VictoriaMetrics cluster mode a tenant is accountID[:projectID], each a 32-bit unsigned integer (projectID defaults to 0):
/insert/<accountID>[:<projectID>]/prometheus/api/v1/write
/select/<accountID>[:<projectID>]/prometheus/api/v1/query
- Tenants are created automatically on the first write;
/admin/tenantson vmselect lists them. - Tenant data is spread evenly across all vmstorage nodes (the tenant is part of the series key, not a separate directory), so resource use depends on total active series, not on the number of tenants.
- Since v1.143.0 the tenant can also come from
AccountID/ProjectIDheaders, and since v1.150.0 this is on by default (-enableMultitenancyViaHeaders=true):/select/prometheus/...without a tenant in the path is valid and defaults to0:0. A proxy that authorizes by path must therefore also override these headers. - Cross-tenant reads and writes use
/select/multitenant/...and/insert/multitenant/...with thevm_account_id/vm_project_idlabels. Exposing these endpoints requires pinningextra_labeland blankingextra_filtersin vmauth, because vmselect ORs client-supplied filters. - Tenant metadata (names, tokens, limits) is not stored in VictoriaMetrics; it lives in the proxy config (vmauth, vmgateway) or an external system.
VictoriaLogs and VictoriaTraces use AccountID and ProjectID headers for multitenancy (default 0:0). For the vmauth configuration that injects these headers, see VictoriaLogs Tenant Isolation.
Sources¶
- VictoriaMetrics single-node docs: Storage, IndexDB
- VictoriaMetrics cluster docs: architecture, replication, multitenancy
- VictoriaLogs FAQ: how VictoriaLogs works
- VictoriaLogs cluster docs
- VictoriaTraces docs and querying
- vmauth docs
- VictoriaMetrics CHANGELOG