Skip to content

Explanation

CockroachDB is a distributed SQL database built for cloud-native resilience. It presents a PostgreSQL-compatible SQL interface. Internally it stores data as one sorted key-value map, split into contiguous ranges that are each replicated across nodes by their own Raft group. Transactions default to SERIALIZABLE isolation (READ COMMITTED is also available since v24.1), and a cluster can survive the loss of a zone or a whole region without manual failover.

See also: Overview, How-to Guides, Reference.


Layer Model

CockroachDB is organized as a stack of five cooperating layers. Each layer talks only to the layer directly above or below it, which keeps concerns separate. Every node runs every layer, so any node can act as a SQL gateway.

Layer Responsibility Key Abstractions
SQL Parse, plan, optimize, and execute SQL pgwire protocol, cost-based optimizer, DistSQL
Transaction ACID guarantees, distributed transactions TxnCoordSender, timestamp cache, write intents, parallel commits
Distribution Route requests to the correct range DistSender, range descriptor cache, meta ranges
Replication Consistent replication through Raft Raft groups, leader leases, Raft log
Storage Durable key-value storage Pebble LSM tree, MVCC keys

The diagram below shows the layers on one node and which component in each layer hands a request down to the next.

graph TB
    Client["SQL client (pgwire, port 26257)"]
    SQL["SQL Layer<br/>(Parser / Optimizer / DistSQL executor)"]
    TXN["Transaction Layer<br/>(TxnCoordSender / Timestamp cache)"]
    DIST["Distribution Layer<br/>(DistSender / Range descriptor cache)"]
    REPL["Replication Layer<br/>(Raft groups / Leader leases)"]
    STORE["Storage Layer<br/>(Pebble engine / MVCC)"]

    Client --> SQL
    SQL --> TXN
    TXN --> DIST
    DIST --> REPL
    REPL --> STORE

    style SQL fill:#e8f4f8,stroke:#2196f3,color:#000
    style TXN fill:#fff3e0,stroke:#ff9800,color:#000
    style DIST fill:#e8f5e9,stroke:#4caf50,color:#000
    style REPL fill:#fce4ec,stroke:#e91e63,color:#000
    style STORE fill:#f3e5f5,stroke:#9c27b0,color:#000

SQL Layer

The SQL layer speaks the PostgreSQL wire protocol (pgwire). This makes CockroachDB compatible with most PostgreSQL drivers and ORMs. Its pipeline is:

  1. Parser: converts SQL text into an abstract syntax tree.
  2. Optimizer: a cost-based optimizer (CBO) that explores equivalent plans using relational-algebra transformation rules and picks the cheapest one. It uses table statistics that are collected automatically. It is locality-aware in multi-region clusters.
  3. Executor: runs the physical plan. For distributed queries (DistSQL), it pushes filtering, aggregation and joins to the nodes that hold the data, which cuts network round-trips.
  4. Encoding: rows and index entries become KV pairs. For example, a primary-key row is stored under a key such as /Table/<table_id>/<index_id>/<pk>.

PostgreSQL Compatibility

CockroachDB supports a large and growing subset of PostgreSQL: window functions, CTEs, JSONB, ARRAY, user-defined functions, PL/pgSQL and stored procedures (since v23.2), triggers (GA in v26.2), row-level security (v25.2) and the pgvector-compatible VECTOR type. PostgreSQL extensions (PostGIS as an extension, TimescaleDB, and others) cannot be loaded. Spatial types are built in instead. Some catalog, locking and DDL behaviors differ, so test ORMs and migrations. See SQL feature support.


Transaction Layer

CockroachDB implements distributed ACID transactions with a combination of techniques:

  • Hybrid Logical Clock (HLC) timestamps: every transaction gets an HLC timestamp. Reads see data as of that timestamp. Nodes must keep clocks within the maximum offset (default 500 ms). A node whose clock drifts too far from the rest of the cluster shuts itself down to protect consistency.
  • Write intents and the transaction record: uncommitted writes are stored as MVCC intents that point to a transaction record. The record lives in the range of the transaction's first write and has a status of PENDING, STAGING, COMMITTED, or ABORTED.
  • Timestamp cache: tracks the latest read timestamp per key span. A write below a cached read timestamp is pushed to a higher timestamp, which prevents read-write anomalies.
  • Transaction pipelining: writes are replicated asynchronously while the client continues. The coordinator only waits for all of them at commit.
  • Parallel Commits: the coordinator writes the transaction record in the STAGING state together with the list of in-flight writes, in parallel with the last writes. Once all of them are replicated, the transaction is implicitly committed. That cuts commit latency from two rounds of consensus to one. Intents are resolved asynchronously afterwards.
  • One-phase commit (1PC): a transaction whose writes all land in a single range commits in one Raft round without creating intents.
  • Read refreshes and retries: if a transaction's timestamp is pushed, it tries to refresh its reads at the new timestamp. Only if the refresh fails does the client see a retryable 40001 error.

Serializable by default, not serializable only

CockroachDB runs transactions at SERIALIZABLE isolation by default. Since v24.1, READ COMMITTED is generally available and enabled by default through sql.txn.read_committed_isolation.enabled. Sessions or transactions opt in with SET TRANSACTION ISOLATION LEVEL READ COMMITTED. CockroachDB's READ COMMITTED is stronger than PostgreSQL's (it prevents anomalies within a single statement) and never returns serialization errors that need client-side retries. READ UNCOMMITTED is upgraded to READ COMMITTED. Cockroach Labs recommends staying on SERIALIZABLE unless contention-driven retries or a migration require otherwise (Read Committed).


Distribution Layer

The distribution layer maps logical key ranges to physical nodes:

  • Range descriptors: the keyspace is divided into ranges. Each range has a descriptor listing its replicas. Descriptors are stored in two-level meta ranges (meta1, meta2).
  • DistSender: the gateway node's DistSender splits a BatchRequest by range and sends each part to that range's leaseholder. It caches range descriptors and refreshes them when it hits a stale entry.
  • Range splits: a range splits automatically when it grows past range_max_bytes (default 512 MiB) or gets a disproportionate share of load (load-based splitting). ALTER TABLE ... SPLIT AT splits manually, for example to isolate a workload. Small, quiet neighbors merge when they are below range_min_bytes (default 128 MiB).

Replication Layer

Each range is replicated by its own Raft group. This layer provides strong consistency and fault tolerance:

  • Replicas: each range has num_replicas replicas (default 3) on different nodes, placed according to zone configurations. Non-voting replicas can serve follower reads without taking part in quorum.
  • Leaseholder: one replica holds the range lease. It serves strongly consistent reads locally, without a Raft round, and proposes all writes.
  • Raft log: an ordered, on-disk log of commands agreed by the voting replicas. A command commits once a majority of voters have appended it.
  • Leader leases (default since v25.2): the lease is tied to Raft leadership through a store-wide store liveness mechanism ("fortification"). The leaseholder is therefore always the Raft leader, except briefly during lease transfers. This removes the old node-liveness range as a single point of failure. It also removes the leader-leaseholder split failure mode of epoch-based leases, where a partitioned leaseholder could keep its lease while unable to propose writes. Meta and system ranges still use expiration-based leases (Replication layer).
  • Per-replica circuit breakers: a replica that cannot make progress for 1 minute (default) trips its breaker and returns ReplicaUnavailableError quickly, instead of letting queries hang.

The diagram below shows how a leaseholder that is also the Raft leader replicates a write to its followers.

graph LR
    GW["Gateway DistSender"]
    LH["Leaseholder + Raft leader<br/>(Node A)"]
    R1["Follower voter<br/>(Node B)"]
    R2["Follower voter<br/>(Node C)"]

    GW -->|"BatchRequest"| LH
    LH -->|"MsgApp (log entry)"| R1
    LH -->|"MsgApp (log entry)"| R2
    R1 -->|"Append ack"| LH
    R2 -->|"Append ack"| LH
    LH -->|"Quorum reached, entry applied"| GW

    style LH fill:#e8f5e9,stroke:#4caf50,color:#000
    style R1 fill:#e3f2fd,stroke:#2196f3,color:#000
    style R2 fill:#e3f2fd,stroke:#2196f3,color:#000

Read Performance

Strongly consistent reads go only to the leaseholder, with no Raft round-trip. A single-range point read costs about one local key-value lookup plus the gateway-to-leaseholder hop. Cockroach Labs quotes about 1 ms for single-row reads within one availability zone.

Zone Configurations

Zone configs control replication topology for a cluster, database, table, index, or partition. They specify the replica count, voter count, placement constraints, lease preferences, range size bounds and MVCC garbage-collection TTL. The variables and their defaults are listed in Reference: Zone Configuration Variables. Multi-region SQL (ALTER DATABASE ... SET PRIMARY REGION, table localities, survival goals) generates zone configs for you and is the recommended interface. Raw zone configs remain available for fine-tuning.


Storage Layer (Pebble)

Pebble is CockroachDB's purpose-built storage engine, written in Go. It is an LSM-tree (Log-Structured Merge-tree) key-value store inspired by LevelDB and RocksDB, with several CockroachDB-specific optimizations. Each node has one or more stores (one Pebble instance per --store).

LSM Tree Structure

The diagram below traces how a write moves through Pebble, from the WAL and memtable down through the compacted levels.

graph TB
    WRITE["Applied Raft command"]
    WAL["Write-Ahead Log<br/>(WAL)"]
    MEM["MemTable<br/>(in-memory skiplist)"]
    L0["L0 SSTables<br/>(sorted, may overlap)"]
    L1["L1 SSTables<br/>(sorted, partitioned)"]
    L2["L2 SSTables<br/>(about 10x L1)"]
    LPLUS["L3 ... L6<br/>(each about 10x prior)"]

    WRITE --> WAL
    WRITE --> MEM
    MEM -->|"Flush when full"| L0
    L0 -->|"Compaction"| L1
    L1 -->|"Compaction"| L2
    L2 -->|"Compaction"| LPLUS

    style MEM fill:#fff3e0,stroke:#ff9800,color:#000
    style WAL fill:#fce4ec,stroke:#e91e63,color:#000
    style L0 fill:#e8f5e9,stroke:#4caf50,color:#000

Key Pebble internals:

Component Role
MemTable In-memory sorted structure (skiplist). All writes land here first.
WAL Sequential write-ahead log for crash recovery of in-flight MemTable data.
L0 SSTables Flushed MemTables. Sorted within each file, but key ranges can overlap across files. Too many L0 files (read amplification) triggers admission control.
L1-L6 SSTables Compacted levels. Each level is partitioned into non-overlapping key ranges and is about 10x the size of the previous one.
Bloom Filters Per-SSTable bloom filters reduce disk reads for point lookups on keys that do not exist.
Block Cache Frequently accessed data blocks are cached in memory to reduce disk I/O.
Compaction Background process that merges SSTables from one level into the next and drops shadowed versions and deleted entries.

MVCC Integration

CockroachDB keys carry MVCC timestamps, so keys look like <key>/<timestamp>. Pebble stores them in sorted order, which keeps all versions of a key next to each other. Versions older than the zone's gc.ttlseconds (default 4 hours) are garbage-collected by the MVCC GC queue, and compactions reclaim the space.

Why Pebble (not RocksDB)

CockroachDB originally used RocksDB. It switched to Pebble (the default since v20.2) to:

  • Remove cgo overhead and the complexity of managing a C++ dependency from Go.
  • Gain tighter control over compaction heuristics, especially MVCC-aware garbage collection.
  • Implement CockroachDB-specific features such as range tombstones for efficient large deletions and ingestion of external SSTables, which IMPORT, RESTORE, and index backfills use.

Data Flow: End-to-End Write Path

The following diagram traces a single-row INSERT sent as an implicit transaction. All of its writes land in one range, so it takes the one-phase-commit (1PC) fast path: one Raft round and no intents.

sequenceDiagram
    participant C as SQL Client
    participant GW as Gateway Node<br/>(SQL + TxnCoordSender)
    participant LH as Leaseholder and Raft leader<br/>(Range R1)
    participant F1 as Follower voter
    participant F2 as Follower voter
    participant P as Pebble (each replica)

    C->>GW: INSERT INTO orders (id, total) VALUES (42, 99.5)
    GW->>GW: Parse, optimize, encode as KV Put<br/>on /Table/orders_id/1/42
    GW->>LH: BatchRequest [ConditionalPut + EndTxn] via DistSender
    LH->>LH: Acquire latches, check timestamp cache,<br/>evaluate as 1PC
    LH->>F1: MsgApp (Raft log entry)
    LH->>F2: MsgApp (Raft log entry)
    F1-->>LH: Append ack (quorum = leader + F1)
    LH->>P: Apply entry, write committed MVCC value
    LH-->>GW: BatchResponse OK
    GW-->>C: INSERT 0 1
    F2-->>LH: Append ack (late, not on the critical path)

Replication Topology Patterns

Single-Region (3 replicas)

Three nodes spread across three availability zones in one region. Survives the loss of one zone. This is the standard HA deployment.

Multi-Region (region survival)

At least three regions with at least three nodes each. With SURVIVE REGION FAILURE, each range gets 5 voters spread so that losing any single region keeps a quorum. Writes pay a cross-region round-trip to reach quorum.

Regional-by-Row and Global Tables

REGIONAL BY ROW tables keep each row's leaseholder and voters in the region named by the hidden crdb_region column, so local reads and writes stay local. GLOBAL tables use non-blocking transactions and global_reads: every region can do fast, consistent reads, and writes pay extra latency. Use them for rarely written reference data.

Latency Trade-offs

A Raft majority must acknowledge every write. In multi-region deployments, choose survival goals (ZONE vs REGION) and table localities deliberately. Region survival means every write crosses a WAN link. Zone survival keeps writes in-region but loses availability if the region fails.


How It Works

Range-based sharding, Raft consensus, the distributed transaction protocol, and leaseholder reads.

Data Distribution

The diagram below shows how the sorted keyspace is cut into ranges and how each range's three replicas and its leaseholder are spread across nodes.

flowchart TB
    subgraph Keyspace["Sorted keyspace"]
        R1["Range 1<br/>/Meta1 - /System"]
        R2["Range 2<br/>/Table/users a-m"]
        R3["Range 3<br/>/Table/users n-z"]
        R4["Range 4<br/>/Table/orders"]
    end

    subgraph Cluster_C["3-node cluster"]
        N1["Node 1<br/>R1 leaseholder, R2, R3, R4 leaseholder"]
        N2["Node 2<br/>R1, R2 leaseholder, R3, R4"]
        N3["Node 3<br/>R1, R2, R3 leaseholder, R4"]
    end

    R1 -.->|"3 replicas"| Cluster_C
    R2 -.->|"3 replicas"| Cluster_C
    R3 -.->|"3 replicas"| Cluster_C
    R4 -.->|"3 replicas"| Cluster_C

    style Keyspace fill:#6933ff,color:#fff

Distributed Transaction (Write)

An explicit transaction that updates two ranges uses pipelined writes and Parallel Commits. The transaction record lives in Range A, where the first write happened. The client gets its COMMIT acknowledgement once the STAGING record and every in-flight intent are replicated. Marking the record COMMITTED and resolving the intents then happens asynchronously.

sequenceDiagram
    participant App as Client (pgwire)
    participant GW as Gateway TxnCoordSender
    participant LHA as Leaseholder Range A (users)
    participant LHB as Leaseholder Range B (orders)
    participant FA as Range A followers

    App->>GW: BEGIN
    App->>GW: UPDATE users SET ... WHERE id = 1
    GW->>LHA: Write intent (pipelined, replication async)
    GW-->>App: UPDATE 1
    App->>GW: UPDATE orders SET ... WHERE id = 7
    GW->>LHB: Write intent (pipelined, replication async)
    GW-->>App: UPDATE 1
    App->>GW: COMMIT
    par Parallel Commits
        GW->>LHA: EndTxn, record STAGING with in-flight writes
        LHA->>FA: Raft replicate STAGING record
        FA-->>LHA: Quorum ack
    and
        GW->>LHB: QueryIntent, confirm orders intent replicated
    end
    GW-->>App: COMMIT ok (implicitly committed)
    GW->>LHA: Async mark record COMMITTED, resolve intent
    GW->>LHB: Async resolve intent to committed value

If the coordinator crashes after STAGING, another transaction that hits one of the intents runs a status recovery. It checks whether every in-flight write succeeded. If they all did, the transaction is committed. Otherwise it is aborted (Transaction layer).

Leaseholder Reads (Fast Path)

Each range has a leaseholder, the replica that serves consistent reads without Raft consensus:

Read Type Mechanism Latency driver
Leaseholder read Read from the leaseholder directly Gateway-to-leaseholder hop, about 1 ms in-zone
Follower read Read from the closest replica at a slightly stale timestamp Local replica, no cross-region hop
Consistent read from a remote region Must go to the leaseholder Adds one cross-region RTT

Follower Reads

CockroachDB supports follower reads: reads served by the nearest replica, including non-voting replicas, instead of the leaseholder:

  • Exact staleness reads (AS OF SYSTEM TIME follower_read_timestamp() or a fixed interval) read from a local replica at a timestamp far enough in the past (a few seconds by default) that the replica has a closed timestamp covering it.
  • Bounded staleness reads (with_max_staleness() / with_min_timestamp()) let CockroachDB pick the freshest timestamp that the local replica can serve within the bound. They must be single-statement (implicit) transactions that read a single row without an index join.

This cuts read latency in geo-distributed deployments by avoiding cross-region round trips to the leaseholder.

Range Splits and Merges

CockroachDB splits and merges ranges automatically:

  • Split trigger: a range exceeds range_max_bytes (512 MiB default), or load-based splitting sees sustained QPS or CPU above threshold (kv.range_split.load_cpu_threshold, 500 ms/s default).
  • Merge trigger: adjacent ranges are below range_min_bytes and not hot.
  • Rebalancing: the allocator and store rebalancer move replicas and leases to balance disk use and load across stores. v26.3 turns on the newer multi-metric allocator (MMA) by default: kv.allocator.load_based_rebalancing now defaults to auto, which switches to multi-metric and count after upgrade finalization to v26.3, so mixed-version clusters keep the legacy rebalancer (kvserverbase/base.go, checked 2026-09-28; issue #169411).

Vector Search Internals

CockroachDB stores embeddings in a pgvector-compatible VECTOR(n) column type (preview in v24.2, GA in v25.4) and indexes them with a vector index (preview in v25.2, GA in v25.4):

  • The index organizes vectors into a hierarchical tree of partitions built by k-means clustering. Similar vectors share a partition. Partitions split when they grow beyond max_partition_size and merge when they fall below min_partition_size.
  • A query descends the tree and explores the vector_search_beam_size closest partitions per level (default 32), then ranks candidates with the opclass's distance metric (L2, cosine, or inner product). That makes it an approximate nearest-neighbor (ANN) search whose recall is tuned by beam size and partition size.
  • Because the index is stored as ordinary KV data in ranges, it is transactionally consistent with the base table, replicated by Raft, and distributed like any other index. There is no separate vector store to keep in sync.
  • Prefix columns (for example (customer_id, embedding)) partition the index by tenant or category, so filtered searches only visit the relevant sub-tree. Filters that do not match a prefix column are not accelerated.

Parameters and limitations are in Reference: Vector Index Parameters. The 25.2 announcement is at Vector indexing blog.


Licensing Model History

  • 2015 to 2019: Apache 2.0 core plus a Cockroach Community License (CCL) for enterprise features.
  • 2019 (v19.2) to 2024: the core moved to the Business Source License (BSL 1.1). BSL code converted to Apache 2.0 after three years. Enterprise features stayed under the CCL and needed a paid key. This two-tier model was known as "CockroachDB Core" vs "Enterprise".
  • 2024-11-18 (v24.3) onward: Cockroach Labs retired Core. All releases, plus later patches of v23.1 to v24.2, ship under the proprietary CockroachDB Software License (CSL) with one feature set, CockroachDB Enterprise. Access is by license key: paid Enterprise, Enterprise Free (organizations under $10M annual revenue, with mandatory telemetry and annual renewal), or a 30-day Trial. Clusters without a valid key, or without required telemetry, are throttled to 5 concurrent transactions. Single-node development clusters need no key (Licensing FAQs).

Why it matters: the source code is still public on GitHub, but CockroachDB is no longer open source in any OSI sense, and production use by larger companies now requires a commercial agreement. Features that used to need a separate enterprise key (for example encryption at rest, enterprise changefeeds, and geo-partitioning) are now available to every licensed cluster. Projects that wanted a permissively licensed distributed SQL database have looked at alternatives. For example, ZITADEL deprecated CockroachDB in v3 in favor of PostgreSQL.


Release Model

Since 2024, Cockroach Labs has shipped a major version every quarter and alternated two release types (Releases overview):

  • Regular releases (for example v25.2, v25.4, v26.2) get 12 months of maintenance plus 6 months of assistance support. After a series proves stable, later patches are designated LTS, which resets the clock to 12 months of maintenance plus 12 months of assistance from the first LTS patch.
  • Innovation releases (for example v25.3, v26.1, v26.3) carry the newest features but get only 6 months of support and no LTS. Self-hosted and CockroachDB Cloud Advanced users may skip them. Cloud Basic and Standard never receive them.

The design trade-off: fast feature delivery for teams that upgrade often, and a slower, required upgrade path (Regular to Regular, about every six months) for teams that value stability. Major-version upgrades are rolling (one node at a time) and are finalized afterwards. Until finalization, a cluster can be rolled back to the previous binary. Support dates are in Reference: Release and Support Matrix.


Benchmarks

The most recent TPC-C figures in the official docs are from v21.1: 1,684,437 tpmC at 140,000 warehouses on 81 nodes (95.45% efficiency), in SERIALIZABLE isolation. Cockroach Labs also reports near-linear KV95 scaling up to 256 nodes, and single-row reads of about 1 ms and writes of about 2 ms within one availability zone. Tables and caveats are in Reference: Benchmarks.

Why distributed SQL is slower per operation

Every write needs a Raft quorum, and cross-range transactions need intents plus a commit protocol. A single-row write is therefore slower than on a single-node PostgreSQL primary with local WAL fsync, especially across zones or regions. CockroachDB's advantage is aggregate throughput, horizontal scale, and automatic failover, not single-operation latency.


Security Model

CockroachDB is secure by default in production mode: cockroach start without --insecure requires TLS certificates. The system layers authentication, authorization, encryption in transit and at rest, and audit logging. Since the Core license was retired (v24.3), every feature below is available to any licensed cluster (paid Enterprise, Enterprise Free, or Trial). Setup commands are in How-to Guides: Secure a Cluster, and the checklist is in Reference: Security Best Practices.

Authentication

CockroachDB supports several client authentication methods. Host-based authentication rules (server.host_based_authentication.configuration, in pg_hba.conf syntax) decide which method applies to which user and source address.

TLS Certificate Authentication

Certificate authentication is the recommended production method. CockroachDB uses a PKI model:

  • A cluster CA signs node certificates and client certificates.
  • Each node presents a node certificate to other nodes (mutual TLS for inter-node traffic).
  • Each SQL client presents a user certificate. The certificate's CN maps to the SQL username. Mapping from the Subject Alternative Name is in preview in v26.2.

Password Authentication

Passwords are hashed with SCRAM-SHA-256 by default (server.user_login.password_encryption = scram-sha-256). Older bcrypt hashes are automatically upgraded. Password authentication is common on CockroachDB Cloud.

Password Auth Requires TLS

Use password authentication only over TLS-encrypted connections. Password authentication without TLS is only possible in special configurations that the docs list as a preview feature. --insecure mode skips authentication entirely and must never be used in production.

GSSAPI / Kerberos Authentication (Enterprise)

SQL clients can authenticate with Kerberos tickets from Active Directory or an MIT KDC. The node needs a keytab (KRB5_KTNAME), and an HBA rule such as host all all all gss include_realm=0 in server.host_based_authentication.configuration selects GSSAPI (GSSAPI authentication).

Cluster Single Sign-On (Enterprise)

JWT-based cluster SSO: users authenticate with an external identity provider (IdP) that issues a JSON Web Token. The cluster validates the JWT signature against configured issuers and JWKS and maps the token subject to a SQL user. Since v25.4 (preview), group claims can automatically sync role memberships, and users can be provisioned automatically. LDAP authentication is also supported (since v24.3).

DB Console SSO (Enterprise)

The DB Console supports OpenID Connect (OIDC) SSO with corporate IdPs such as Okta, Microsoft Entra ID, or Google Workspace. Just-in-time provisioning of OIDC users is in preview in v26.1.

Authorization (RBAC)

CockroachDB implements role-based access control modeled on PostgreSQL's privilege system:

  • Users are roles that can log in. Every connection belongs to a user.
  • Roles group privileges. Users inherit privileges from roles they are granted.
  • The admin role has full cluster access, and root is a member of admin. Since v26.1 (preview), root can be disabled.
  • Row-level security (v25.2) adds per-row policies (CREATE POLICY), for example for multi-tenant tables or for restricting what AI agents can read.
  • ALTER DEFAULT PRIVILEGES sets privileges for objects created in the future.

Privilege levels are listed in Reference: Privilege Hierarchy.

Encryption at Rest (Enterprise)

Encryption at rest works at the Pebble level. It encrypts all store files (SSTables, WAL, manifests) with AES in counter mode (AES-CTR), using 128-, 192-, or 256-bit keys:

  • A user-supplied store key (a key file generated with cockroach gen encryption-key) encrypts a registry of data keys.
  • Data keys encrypt the actual files and rotate automatically every rotation-period (default one week) or whenever the store key changes.
  • The store key is rotated by restarting a node with key=<new> and old-key=<previous> in the --enterprise-encryption flag.
  • On self-hosted clusters, keys are files on the node. Customer-managed keys in a cloud KMS (CMEK) are a CockroachDB Cloud Advanced feature, not a cockroach start option.

Correction

Earlier versions of this note showed aws-kms:// URIs in --enterprise-encryption and an enterprise_key zone-config variable. Neither is documented. Self-hosted encryption at rest uses key files per store (Encryption at rest).

Encryption in Transit

All inter-node traffic (Raft, gossip, DistSQL flows) and client traffic uses TLS in secure mode. Nodes authenticate each other with mutual TLS. Clients verify the node certificate and can present their own certificate. --tls-cipher-suites (v25.2) restricts allowed suites. For TLS 1.3, v26.2 uses the hybrid post-quantum key exchange X25519MLKEM768 by default (preview) for both client-to-node and inter-node connections. Certificate revocation is available through OCSP. CRLs are not supported.

Audit Logging

  • Role-based audit logging: the sql.log.user_audit cluster setting lists roles and statement types to audit. Events are emitted as role_based_audit_event on the SENSITIVE_ACCESS log channel.
  • Table-based audit logging: ALTER TABLE ... EXPERIMENTAL_AUDIT SET READ WRITE logs every query touching a table.
  • System event log: all clusters record DDL, privilege changes, user creation, zone-config changes and node join/leave events in system.eventlog and in the structured log channels.

The feature-availability page (v26.2) still lists both audit-logging modes under preview features.

Network Security

  • Host-based authentication rules restrict which users can connect from which CIDR ranges and with which method.
  • CockroachDB Cloud adds AWS PrivateLink, GCP Private Service Connect / VPC peering, Azure Private Link, and IP allowlists. It also offers egress private endpoints (limited access in v26.2).

Sources