Skip to content

Explanation

What this page covers

How the Victoria Stack works and why it is built that way: topology, the shared insert/select/storage pattern, storage engines (and why there is no WAL), the write and read paths, VictoriaTraces on top of the VictoriaLogs engine, the security and multitenancy model, and the open-core split. For exact flags, ports and versions see Reference; for recipes see How-to guides; hub: Victoria Stack.

Default Topology

The reference production topology puts collectors in front of a single vmauth entry point that routes by URL path to three independent storage systems, with Grafana and vmalert as consumers:

flowchart TB
    subgraph Sources["Data Sources"]
        K8s["Kubernetes<br/>Pods and Services"]
        Apps["Applications<br/>(OTel SDK)"]
        Infra["Exporters<br/>(node_exporter, etc.)"]
        Logs["Log shippers<br/>(Fluent Bit, vlagent, Vector)"]
    end

    subgraph Collection["Collection Layer"]
        Agent["vmagent<br/>scrape + remote write"]
        OTel["OTel Collector<br/>(optional)"]
    end

    subgraph Proxy["Routing Layer"]
        Auth["vmauth<br/>auth, route, load balance"]
    end

    subgraph MetricsCluster["VictoriaMetrics (metrics)"]
        MI["vminsert x2"]
        MS["vmstorage x3<br/>(StatefulSet, SSD)"]
        MSel["vmselect x2"]
        MI --> MS
        MSel --> MS
    end

    subgraph LogsCluster["VictoriaLogs (logs)"]
        LI["vlinsert x2"]
        LS["vlstorage x3"]
        LSel["vlselect x2"]
        LI --> LS
        LSel --> LS
    end

    subgraph TracesCluster["VictoriaTraces (traces)"]
        TI["vtinsert x2"]
        TS["vtstorage x3"]
        TSel["vtselect x2"]
        TI --> TS
        TSel --> TS
    end

    subgraph Alerting["Alerting"]
        Alert["vmalert"]
        AM["Alertmanager"]
    end

    subgraph Viz["Visualization"]
        Grafana["Grafana"]
        VMUI["VMUI<br/>(built-in)"]
    end

    K8s --> Agent
    Infra --> Agent
    Apps --> OTel
    Agent --> Auth
    OTel --> Auth
    Logs --> Auth

    Auth -->|"/api/v1/write"| MI
    Auth -->|"/insert/jsonline, /insert/loki"| LI
    Auth -->|"/insert/opentelemetry/v1/traces"| TI

    Auth -->|"PromQL / MetricsQL"| MSel
    Auth -->|"LogsQL"| LSel
    Auth -->|"/select/jaeger, /select/tempo"| TSel

    Grafana --> Auth
    VMUI --> Auth
    Alert --> Auth
    Alert -->|"fire alerts"| AM

    style Sources fill:#0d7377,color:#fff
    style Collection fill:#ff6600,color:#fff
    style Proxy fill:#7b42bc,color:#fff
    style MetricsCluster fill:#2a2d3e,color:#fff
    style LogsCluster fill:#2a7de1,color:#fff
    style TracesCluster fill:#e65100,color:#fff
    style Alerting fill:#c62828,color:#fff
    style Viz fill:#ff6600,color:#fff

The three databases share design ideas and tooling (vmauth, vmalert, VMUI, operator) but not storage: each has its own storage nodes, retention and backups.

Deployment Modes

Single-Node vs Cluster

Feature Single-Node Cluster
Scalability Vertical only Horizontal and vertical
Operational Complexity Very low (1 binary) Moderate (3 roles)
Multi-tenancy No (VictoriaMetrics); yes for VictoriaLogs/VictoriaTraces via headers Yes (accountID:projectID)
Replication No (relies on durable disk) VictoriaMetrics: -replicationFactor=N; VictoriaLogs/VictoriaTraces: none (shard only)
Target Workload Upstream recommends single-node below ~1M samples/s Beyond one machine
External Dependencies None None

Recommendation (upstream): start with single-node. Move to cluster only for multitenancy, horizontal scale beyond one machine, or application-level replication. A single-node VictoriaLogs or VictoriaTraces instance can later be listed as a storage node of a cluster, so the migration path is incremental.

Component Roles

Each signal follows the same tri-role pattern in cluster mode:

Role Metrics Logs Traces
Ingestion (stateless) vminsert vlinsert vtinsert
Querying (stateless) vmselect vlselect vtselect
Storage (stateful) vmstorage vlstorage vtstorage
Binaries Three separate binaries One victoria-logs binary; role set by -storageNode, -insert.disable, -select.disable One victoria-traces binary, same model
Sharding key Consistent hash of metric name + labels Even spread of log streams Trace ID

Insert and select roles are stateless and scale on CPU; storage nodes hold the data and scale on disk and memory. vmstorage nodes never talk to each other (shared-nothing), which removes the need for a consensus protocol.

Core Mechanisms

Storage Engine (VictoriaMetrics)

  1. Buffer, then immutable parts. Incoming samples are buffered in memory for about a second, turned into searchable in-memory parts, and periodically persisted (-inmemoryDataFlushInterval) to part directories under <storageDataPath>/data/small/YYYY_MM/. Background merges combine small parts into bigger ones (data/big/), an LSM-like design.
  2. No write-ahead log. VictoriaMetrics deliberately has no WAL. On an unclean shutdown (OOM, kill -9, power loss) the last few seconds of unflushed data can be lost. Upstream's argument: senders (Prometheus, vmagent with its on-disk queue) already retry, and a WAL costs disk IO on every write. See this article.
  3. Blocks by TSID. Each part stores blocks of up to 8K samples sorted by internal TSID; timestamps and values are encoded separately and compressed (ZSTD plus type-specific encodings).
  4. Per-partition IndexDB. The inverted index (label pairs -> TSID) has global and per-day variants. Since v1.133.0 (2026-01-02) every monthly partition has its own IndexDB, which is dropped with the partition when it leaves retention. The upgrade re-registers all active series once.
  5. Deterministic sharding. In cluster mode vminsert uses consistent hashing over metric name and labels to pick vmstorage nodes. Since v1.149.0 slowness-based rerouting is on by default: if one vmstorage node is slow, new series are routed elsewhere (disabled automatically when -replicationFactor > 1).

VictoriaMetrics Write and Query Internals

A compact view of how the single-node engine connects ingestion APIs, storage and the query engine:

flowchart LR
    subgraph Ingestion["Ingestion APIs"]
        PR["Prometheus<br/>remote_write"]
        IL["InfluxDB<br/>line protocol"]
        DD["Datadog / NewRelic"]
        OT["OTLP metrics"]
        GR["Graphite"]
    end

    subgraph VM["VictoriaMetrics storage"]
        WB["In-memory buffer<br/>(about 1s)"]
        IMP["In-memory parts<br/>(searchable)"]
        SMALL["data/small/YYYY_MM<br/>parts"]
        BIG["data/big/YYYY_MM<br/>merged parts"]
        IDX["Per-partition IndexDB<br/>(labels to TSID)"]
    end

    subgraph Query["Query"]
        MQL["MetricsQL engine"]
    end

    Ingestion --> WB --> IMP --> SMALL -->|"background merge"| BIG
    WB --> IDX
    MQL --> IDX
    MQL --> IMP
    MQL --> SMALL
    MQL --> BIG
    MQL --> Grafana["Grafana / VMUI"]

    style VM fill:#2a2d3e,color:#fff
    style Ingestion fill:#0d7377,color:#fff
    style Query fill:#ff6600,color:#fff

VictoriaLogs (Logs)

VictoriaLogs accepts logs as JSON-like entries with _msg, _time and _stream plus arbitrary fields:

flowchart LR
    subgraph Ingestion["Ingestion APIs"]
        LK["Loki push"]
        ES["Elasticsearch bulk"]
        SL["Syslog"]
        OT["OTLP logs"]
        JL["JSON lines"]
    end

    subgraph VL["VictoriaLogs storage"]
        direction TB
        Parse["Parse fields<br/>and _stream"]
        Blocks["Per-field column blocks<br/>grouped by stream and time"]
        BF["Bloom filters per block<br/>(word and phrase skip)"]
        Parts["partitions/YYYYMMDD<br/>(per-day)"]
    end

    Ingestion --> Parse --> Blocks --> Parts
    Blocks --> BF

    LogsQL["LogsQL engine<br/>(parallel per CPU core)"] --> BF
    LogsQL --> Parts
    LogsQL --> Grafana["Grafana plugin / VMUI"]

    style VL fill:#2a7de1,color:#fff
    style Ingestion fill:#0d7377,color:#fff
  • Columnar by field. Values of the same field across entries are stored in one block, so queries read only the fields they need. The design is inspired by ClickHouse.
  • Bloom filters, not an inverted index. Word and phrase filters consult per-block bloom filters to skip blocks that cannot match, then scan the rest in parallel on all CPU cores. This keeps RAM low; the trade-off is scan CPU for unselective queries.
  • Stream locality. Entries with the same _stream are stored together, which improves compression and lets stream filters ({app="nginx"}) skip unrelated blocks. A sparse index on _time speeds time filters.
  • Per-day partitions at <storageDataPath>/partitions/YYYYMMDD make retention a directory delete and allow attach/detach of partitions and per-partition snapshots for backup.

VictoriaTraces (Traces)

VictoriaTraces converts OTLP spans into structured log entries and stores them with the VictoriaLogs engine:

flowchart LR
    subgraph Ingestion["Ingestion (OTLP only)"]
        OTLP_H["OTLP/HTTP<br/>/insert/opentelemetry/v1/traces<br/>:10428"]
        OTLP_G["OTLP/gRPC<br/>:4317 when<br/>-otlpGRPCListenAddr set"]
    end

    subgraph VT["VictoriaTraces"]
        direction TB
        TP["Span to structured<br/>log entry"]
        VLS["VictoriaLogs<br/>storage engine"]
        SG["Service graph<br/>relations"]
    end

    Ingestion --> TP --> VLS
    TP --> SG

    JQ["Jaeger query API<br/>/select/jaeger"] --> VLS
    TQ["Tempo API (experimental)<br/>/select/tempo"] --> VLS
    LQ["LogsQL<br/>/select/logsql"] --> VLS
    JQ --> Grafana["Grafana<br/>(Jaeger or Tempo data source)"]
    TQ --> Grafana

    style VT fill:#e65100,color:#fff
    style Ingestion fill:#0d7377,color:#fff
  • Inherits the VictoriaLogs engine (columnar blocks, bloom filters, per-day partitions, compression) and therefore needs no object storage.
  • Ingestion is OTLP only. Jaeger/Zipkin clients should export through the OpenTelemetry Collector. gRPC is off by default and uses TLS unless -otlpGRPC.tls=false.
  • Query surfaces. Jaeger query JSON API (Grafana Jaeger data source, Jaeger UI), LogsQL over spans, and an experimental subset of the Grafana Tempo HTTP API under /select/tempo: search with TraceQL, tag autocomplete, trace by ID (v1 and v2), and /api/metrics/query_range for TraceQL metrics. Tempo support started in v0.8.0 (2026-03-02), was extended for Grafana Traces Drilldown in v0.9.0 (2026-05-20), and upstream lists v0.9.4 as the minimum version for it. Some TraceQL functions and drilldown panels are still unsupported.
  • Maturity. VictoriaTraces is still 0.x (v0.11.1, 2026-09-16). Upstream claims up to 3.7x less RAM and 2.6x less CPU than Tempo, but there is no independent public benchmark at 100M+ spans/day.

Data Flow

The cluster write and read paths for metrics (logs and traces follow the same shape with their own roles):

sequenceDiagram
    participant App as Targets / Apps
    participant Agent as vmagent
    participant Auth as vmauth
    participant Ins as vminsert
    participant Sto as vmstorage nodes
    participant Sel as vmselect
    participant G as Grafana

    Agent->>App: Scrape /metrics
    Agent->>Agent: Relabel, buffer to on-disk queue if remote is down
    Agent->>Auth: Remote write (Prometheus RW)
    Auth->>Auth: Authenticate, pick url_map route
    Auth->>Ins: /insert/0/prometheus/api/v1/write
    Ins->>Sto: Consistent hash per series, RF copies
    Note over Sto: Buffer about 1s, in-memory parts, flush, background merge
    G->>Auth: PromQL or MetricsQL query
    Auth->>Sel: /select/0/prometheus/api/v1/query_range
    Sel->>Sto: Fetch blocks for matching TSIDs from all nodes
    Sto-->>Sel: Compressed blocks
    Sel->>Sel: Deduplicate, merge, evaluate MetricsQL
    Sel-->>G: Result (may be marked partial)
  1. vmagent scrapes targets (or receives pushes), applies relabeling and stream aggregation, and buffers to -remoteWrite.tmpDataPath if the destination is unavailable.
  2. vmauth authenticates the request and routes it by path, host, header or JWT claim to the right backend and tenant.
  3. vminsert shards series across vmstorage nodes (N copies with replication).
  4. vmstorage buffers, flushes and merges as described above; there is no WAL replay on restart.
  5. vmselect queries all vmstorage nodes, deduplicates replicated samples (-dedup.minScrapeInterval), and evaluates MetricsQL. If some nodes are down it returns partial results unless fewer than -replicationFactor nodes are missing.

vmalert Evaluation Flow

vmalert periodically evaluates rule groups against a datasource (VictoriaMetrics, VictoriaLogs or VictoriaTraces) and writes results back:

sequenceDiagram
    participant A as vmalert
    participant P as vmauth
    participant DS as VictoriaMetrics or VictoriaLogs
    participant RW as Remote write target
    participant AM as Alertmanager

    Note over A: Evaluate rule group every interval
    A->>P: Instant query (MetricsQL or LogsQL stats)
    P->>DS: Route by path
    DS-->>P: Result
    P-->>A: Result
    alt Alerting rule fires
        A->>AM: Send alert notification
        A->>RW: Write ALERTS and ALERTS_FOR_STATE series
    else Recording rule
        A->>RW: Remote write recorded series
    end

Alert state is restored after restart from the ALERTS_FOR_STATE series (-remoteRead.url); since tip after v1.152.0 alerts that can be restored straight to firing skip the pending state.

Multi-Source Log Ingestion

VictoriaLogs accepts common shipper protocols directly, so no translation tier is required:

flowchart LR
    A["Promtail / Alloy"] -->|"Loki push API"| B{"vmauth"}
    C["Fluent Bit"] -->|"JSON lines"| B
    D["Logstash / Vector"] -->|"Elasticsearch bulk"| B
    E["OTel Collector"] -->|"OTLP logs"| B
    F["rsyslog"] -->|"Syslog"| G
    V["vlagent"] -->|"native, replicated"| B
    B -->|"route and set tenant headers"| G["vlinsert"]
    G --> H[("vlstorage")]

    style B fill:#7b42bc,color:#fff
    style H fill:#2a7de1,color:#fff

Syslog is a TCP/UDP listener on VictoriaLogs itself (not HTTP), so it bypasses vmauth.

Storage Layout

VictoriaMetrics (Metrics)

<storageDataPath>/
├── data/
│   ├── small/YYYY_MM/     # freshly flushed parts, parts.json per partition
│   └── big/YYYY_MM/       # merged large parts
├── indexdb/               # inverted index (per-partition since v1.133.0)
├── snapshots/             # instant snapshots used by vmbackup
└── metadata/, cache/      # internal state and persisted caches

The metadata/ and cache/ entries and the exact on-disk location of per-partition IndexDB after v1.133.0 are not spelled out in the public storage docs; treat that part of the tree as indicative.

VictoriaLogs (Logs)

<storageDataPath>/
└── partitions/
    └── YYYYMMDD/          # one directory per day; snapshot, attach, detach, delete as a unit

VictoriaTraces (Traces)

Same layout as VictoriaLogs (per-day partitions); spans are organized by trace ID and span attributes.

Key Design Decisions

Decision Rationale Trade-off
No external dependencies No PostgreSQL, Redis, ZooKeeper or object storage — small operational surface Durability depends on the disks and on backups
Local disk over object storage Lower latency, no request costs Capacity bound to node disks; DR needs vmbackup/rsync
No WAL Less write IO; senders already retry Seconds of data can be lost on crash
Shared-nothing cluster Storage nodes do not coordinate — simple scaling Rebalancing is not automatic; old data stays where it was written
Consistent hashing, no consensus Deterministic placement without Raft/Paxos Replication is optional and multiplies cost by RF
Bloom filters (logs, traces) Much less RAM than inverted indexes More CPU for unselective full-text queries
Same binary, multiple roles (logs, traces) A single node can become a storage node of a cluster Role misconfiguration is possible; disable unused APIs
Apache 2.0 + Enterprise features Permissive core; revenue from downsampling, security and support features Some production features (downsampling, mTLS, LTS) are paid

Open-Core Model

All three databases and the vm* tools are Apache 2.0. Enterprise builds (same code base plus closed features, -enterprise image tags) add storage-cost features (downsampling, retention filters), operational automation (vmbackupmanager, automatic vmstorage discovery), security features (vmgateway, mTLS, IP filters, auto-TLS, FIPS builds), vmanomaly, and LTS release lines. Since early 2026 several auth features that previously pushed users towards vmgateway landed in community vmauth: JWT verification (v1.137.0), OIDC discovery and claim matching (v1.138.0), and browser SSO (in tip after v1.152.0). The full matrix is in Reference: Community vs Enterprise.

Security Architecture Overview

VictoriaMetrics components have no built-in per-user authorization. vminsert, vmselect and vmstorage must run in a private network, and all external access goes through an auth proxy: vmauth (community) or vmgateway (Enterprise).

flowchart TD
    subgraph External["External"]
        Clients["API clients<br/>Grafana, vmagent"]
        Users["Browser users"]
    end

    subgraph AuthLayer["Auth proxy layer"]
        Vmauth["vmauth<br/>Basic, Bearer, JWT/OIDC,<br/>mTLS routing (Enterprise)"]
        Vmgateway["vmgateway (Enterprise)<br/>JWT vm_access claims,<br/>rate limits"]
    end

    subgraph Cluster["VictoriaMetrics cluster (private network)"]
        Vminsert["vminsert :8480"]
        Vmselect["vmselect :8481"]
        Vmstorage["vmstorage :8482"]
    end

    Clients -->|"Bearer token or Basic auth"| Vmauth
    Users -->|"OIDC ID token or SSO cookie"| Vmauth
    Users -->|"OIDC JWT"| Vmgateway
    Vmauth -->|"route to tenant path or headers"| Vminsert
    Vmauth -->|"route to tenant path or headers"| Vmselect
    Vmgateway -->|"tenant from vm_access claim"| Vminsert
    Vmgateway -->|"tenant from vm_access claim"| Vmselect
    Vminsert -->|":8400"| Vmstorage
    Vmselect -->|":8401"| Vmstorage

The JWT, OIDC discovery and match_claims capabilities in community vmauth overlap with what used to require vmgateway; vmgateway still adds rate limiting and is Enterprise-only.

Network Topology

Component Network Zone Exposure
vmstorage / vlstorage / vtstorage Private subnet No external access
vminsert / vlinsert / vtinsert Private subnet Behind vmauth only
vmselect / vlselect / vtselect Private subnet Behind vmauth only
vmauth DMZ / public subnet TLS termination point

Multi-Tenant Data Isolation

In VictoriaMetrics cluster mode a tenant is accountID[:projectID], each a 32-bit unsigned integer (projectID defaults to 0):

/insert/<accountID>[:<projectID>]/prometheus/api/v1/write
/select/<accountID>[:<projectID>]/prometheus/api/v1/query
  • Tenants are created automatically on the first write; /admin/tenants on vmselect lists them.
  • Tenant data is spread evenly across all vmstorage nodes (the tenant is part of the series key, not a separate directory), so resource use depends on total active series, not on the number of tenants.
  • Since v1.143.0 the tenant can also come from AccountID/ProjectID headers, and since v1.150.0 this is on by default (-enableMultitenancyViaHeaders=true): /select/prometheus/... without a tenant in the path is valid and defaults to 0:0. A proxy that authorizes by path must therefore also override these headers.
  • Cross-tenant reads and writes use /select/multitenant/... and /insert/multitenant/... with the vm_account_id / vm_project_id labels. Exposing these endpoints requires pinning extra_label and blanking extra_filters in vmauth, because vmselect ORs client-supplied filters.
  • Tenant metadata (names, tokens, limits) is not stored in VictoriaMetrics; it lives in the proxy config (vmauth, vmgateway) or an external system.

VictoriaLogs and VictoriaTraces use AccountID and ProjectID headers for multitenancy (default 0:0). For the vmauth configuration that injects these headers, see VictoriaLogs Tenant Isolation.

Sources