Explanation¶
How NATS works inside: the single-binary server, subject routing, request-reply, cluster routes, superclusters, leaf nodes, the JetStream persistence layer (meta-layer Raft, per-asset Raft, file store), the features added in 2.11 to 2.15, the decentralized security model, and the 2025 governance episode. Look-up tables (config keys, defaults, headers, advisories) live in Reference; tasks live in How-to Guides.
Architecture¶
Component Map¶
Every role runs inside one nats-server process. The diagram shows the listeners, the per-account subject routing core, and the optional JetStream subsystem.
flowchart LR
subgraph Server["nats-server (single Go binary)"]
direction TB
ClientConn["Client listener<br/>(:4222, TLS)"]
WSBridge["WebSocket listener"]
MQTTBridge["MQTT 3.1.1 listener"]
SubsRouter["Subject router<br/>(per-account sublist)"]
AccountIso["Accounts<br/>(exports / imports)"]
ClusterRoute["Cluster routes<br/>(:6222, pooled)"]
Gateway["Gateway endpoint<br/>(:7222 by convention)"]
LeafEndpoint["Leaf node endpoint<br/>(:7422)"]
subgraph JetStreamSubsystem["JetStream subsystem (optional)"]
MetaRaft["Meta-layer Raft group<br/>($JS.API)"]
StreamRaft["Stream Raft groups<br/>(R1 / R3 / R5)"]
ConsumerRaft["Consumer Raft groups"]
FileStore["File store<br/>(message blocks)"]
MemStore["Memory store"]
KVStore["KV bucket = stream KV_name"]
ObjectStore["Object Store = stream OBJ_name"]
end
Monitoring["HTTP monitoring<br/>(:8222 varz, jsz, healthz)"]
end
ClientConn --> SubsRouter
WSBridge --> SubsRouter
MQTTBridge --> SubsRouter
SubsRouter --> AccountIso
AccountIso --> JetStreamSubsystem
AccountIso --> ClusterRoute
AccountIso --> Gateway
AccountIso --> LeafEndpoint
MetaRaft --> StreamRaft
StreamRaft --> FileStore
StreamRaft --> MemStore
StreamRaft --> ConsumerRaft
KVStore --> StreamRaft
ObjectStore --> StreamRaft
Components¶
| Component | Role |
|---|---|
| nats-server | Single static Go binary that hosts every role: client listener, route peer, gateway, leaf node endpoint, JetStream node, MQTT and WebSocket bridges, monitoring server. |
| Subject router | Per-account subscription list ("sublist") that matches published subjects against subscriptions, with * (one token) and > (remaining tokens) wildcards and queue groups. |
| Account | Hard isolation boundary. Subjects, streams, KV buckets and consumers live inside an account. Cross-account traffic only flows through explicit exports and imports. |
System account ($SYS) |
Carries server events, monitoring requests ($SYS.REQ.>), account JWT updates, and by default JetStream Raft replication traffic. |
| Cluster route | Full-mesh TCP links between servers of one cluster that propagate subscription interest. Since 2.10 each peer pair uses a pool of connections (default 3) and accounts can be pinned to dedicated routes. |
| Gateway | Cluster-to-cluster link that forms a supercluster. Interest is propagated per account in "optimistic" or "interest-only" mode (see below). |
| Leaf node | Outbound connection from a remote nats-server into a hub. Leaf links bind local accounts to hub accounts, so the remote side only sees what its account permits. |
| JetStream meta-layer | Cluster-wide Raft group that stores stream and consumer assignments and serves the $JS.API management API. In 2.15 it gained a desired-state reconciliation engine. |
| Stream Raft group | One Raft group per replicated stream (R3 or R5). NATS uses its own Raft implementation (server/raft.go), not a third-party library. |
| Consumer Raft group | Replicated consumers keep delivery and ack state in their own Raft group, placed on the stream's peers. |
| File store | Per-stream append-only message blocks (up to 8 MB each) with per-block indexes, optional S2 compression and optional at-rest encryption. |
| MQTT bridge | MQTT 3.1.1 (QoS 0, 1, 2) implemented on top of JetStream, which stores sessions and retained messages. Requires JetStream. |
Subjects and Wildcards¶
NATS subjects are dot-delimited tokens (orders.created.us-east.123). Two wildcards apply to subscriptions (and to stream subject filters):
*matches exactly one token (orders.*.us-east.123).>matches one or more trailing tokens (orders.>).
The interest graph is account-local, so orders.> in account A never receives messages published in account B unless B exports and A imports that subject. Subscribers in the same queue group share the load: each message goes to one member of the group.
The diagram shows one publish fanning out to two plain subscribers and to one member of a queue group.
flowchart LR
Pub["Publisher<br/>orders.created.us-east.123"]
SubAll["SUB orders.>"]
SubRegion["SUB orders.*.us-east.>"]
Q1["Queue group workers<br/>SUB orders.created.>"]
Q2["Queue group workers<br/>SUB orders.created.>"]
Pub --> SubAll
Pub --> SubRegion
Pub -.->|"one of the group"| Q1
Pub -.->|"one of the group"| Q2
Subject design matters because subjects are also the unit of permissions, stream binding, KV keys, and routing. Core NATS never stores messages: if no subscriber has interest, the message is dropped (at-most-once).
Request-Reply¶
Request-reply is built from plain pub/sub. The client subscribes to a unique inbox subject, publishes the request with that inbox as the reply subject, and waits for the first response. Modern clients multiplex all requests over one wildcard inbox subscription.
sequenceDiagram
participant C as Client
participant S as nats-server
participant Svc as Service (queue group)
C->>S: SUB _INBOX.abc.*
C->>S: PUB svc.req reply _INBOX.abc.1
S->>Svc: MSG svc.req reply _INBOX.abc.1
Svc->>S: PUB _INBOX.abc.1 (response)
S->>C: MSG _INBOX.abc.1
If no responder exists, servers send a "no responders" status so the client fails fast instead of timing out. Placing services in a queue group gives load balancing and failover without a separate load balancer. The NATS docs' single-machine nats bench example averages 50.87 µs per request-reply round trip (Performance Characteristics); network hops add to that.
Cluster Routing¶
Cluster peers form a full mesh of routes. Each server tells its peers which subjects its local clients are interested in (RS+ / RS- protocol operations), so a publishing server knows exactly which peers need a copy. A message crosses at most one route hop inside a cluster.
Since 2.10 routes are pooled (pool_size, default 3) to reduce head-of-line blocking, and specific accounts can be pinned to dedicated route connections. Route traffic can be S2-compressed. Since 2.11 an account can move its JetStream Raft replication traffic out of the system account (cluster_traffic), which helps busy multi-tenant clusters.
Superclusters (Gateways)¶
Gateways connect clusters without a full subscription mesh. For each account, a gateway starts in optimistic mode: it forwards messages and the remote side replies with "no interest" for subjects nobody subscribes to, which the sender caches. When those negative entries grow too large (defaultGatewayMaxRUnsubBeforeSwitch is 1,000 in the source), the account switches to interest-only mode, where the remote cluster sends explicit subscription interest, like a route.
This keeps cross-cluster traffic proportional to real demand. Request-reply across gateways uses reply-subject mapping so responses return to the right cluster. Queue groups prefer local members and only fail over to remote clusters when none are available locally.
Leaf Nodes¶
Leaf nodes connect outbound to a hub cluster or supercluster, which suits firewalled sites and edge devices:
- Each remote binds a local account to a hub account using credentials, so the leaf's traffic lands in exactly one hub account.
- The hub only sees subjects the leaf account is allowed to publish or subscribe to.
- A leaf can run its own JetStream domain. Local streams keep working while the link is down, and streams can mirror or source across domains when it returns.
Since 2.12, isolate_leafnode_interest stops east-west interest propagation between many leaf nodes that never talk to each other. Since 2.14 leaf remotes can be added or removed with a config reload.
Common patterns:
- Edge site - a leaf on a gateway device, local devices speak MQTT or NATS to it, selected subjects flow to the hub.
- Per-tenant leaf - a SaaS customer runs a leaf on-premises and only exports summary subjects.
JetStream¶
JetStream is the persistence layer built into nats-server since 2.2. It replaced the separate NATS Streaming (STAN) server. A stream captures messages published on its subjects. A consumer is a stateful, server-side view of a stream that tracks delivery and acknowledgements.
Storage Model¶
| Aspect | Detail |
|---|---|
| Message blocks | Append-only files per stream. Block size is chosen from stream limits (1, 4 or 8 MB; KV uses 4 MB). |
| Indexes | Per-block subject and sequence indexes, plus a stream state file. State is rebuilt from blocks if the state file is missing or stale. |
| Limits | Messages are removed by max_age, max_bytes, max_msgs, max_msgs_per_subject, per-message TTL (2.11), or by retention policy. |
| Durability | Writes go to the OS page cache and are fsynced every sync_interval (default 2 min). Replicated streams rely on Raft quorum for durability; sync_interval: always fsyncs every write. |
| Compression and encryption | Optional S2 compression per stream (2.10) and per-server at-rest encryption (chachapoly or aes). |
| Memory store | Same API, volatile. Since 2.12 in-memory replicated streams recover more reliably after restarts. |
Retention Policies¶
| Policy | Behaviour |
|---|---|
limits (default) |
Keep messages until a limit is reached. Consumers do not affect retention. |
interest |
Keep a message until every consumer that existed when it arrived has acknowledged it. |
workqueue |
Remove a message once any consumer acknowledges it. Consumers must have non-overlapping filters. |
Consumer Styles¶
| Style | Notes |
|---|---|
| Pull (recommended) | Clients request batches (Fetch, Consume). Scales horizontally; supports priority groups (2.11) with overflow, pinned_client and prioritized (2.12) policies. |
| Push | Server delivers to a deliver_subject; a deliver_group spreads messages over a queue group. |
| Durable vs ephemeral | Durable consumers persist state; ephemeral ones are removed after inactive_threshold. |
| Ordered | Client-managed, single-threaded, ephemeral consumer that recreates itself on gaps. In the newer client APIs (for example the Go jetstream package) it is pull-based. |
Acknowledgement policies are none, all, explicit, and (2.14) flow_control. Replay policies are instant and original.
Stream Replication¶
A replicated stream is a Raft group whose leader orders writes. The sequence shows an R3 publish: the leader replicates the entry to its followers and acknowledges the publisher once a quorum has stored it.
sequenceDiagram
participant P as Publisher
participant L as Stream leader (n1)
participant F1 as Follower (n2)
participant F2 as Follower (n3)
P->>L: PUB orders.created (Nats-Msg-Id)
L->>L: dedup check, append to Raft WAL
L->>F1: AppendEntry
L->>F2: AppendEntry
F1-->>L: ack
Note right of L: quorum 2 of 3 reached
L->>L: apply to file store
L-->>P: PubAck (stream, seq)
F2-->>L: ack (late)
Quorum is a majority (2 of 3, 3 of 5). Since 2.11 a new leader only serves reads and writes after catching up with its Raft log, and deletes in interest and workqueue streams are replicated as proposals. Since 2.12 replicated streams flush to the store asynchronously (the Raft WAL is still written before commit). Since 2.14 leaders that fall behind step down under Raft overrun protection, and filestore I/O errors freeze the affected stream and fail the health check instead of being ignored.
Delivery Guarantees¶
- Core NATS is at-most-once.
- JetStream is at-least-once by default.
- Exactly-once is achieved by combining publish deduplication (
Nats-Msg-Idwithinduplicate_window) with double acknowledgement on the consumer side (AckSync). Redelivery after a crash or ack timeout is still possible if the application does not use both.
Key-Value Store¶
A KV bucket is a stream named KV_<bucket> on subjects $KV.<bucket>.>, with max_msgs_per_subject equal to the configured history (1 by default, up to 64) and direct get enabled. Put appends, Get reads the last message for the key subject, Delete and Purge write markers, Watch is an ordered consumer, and Update (compare-and-swap) uses the Nats-Expected-Last-Subject-Sequence header. Since 2.10 buckets can mirror or source other buckets.
Object Store¶
An Object Store bucket is one stream, OBJ_<bucket>, with two subject spaces: $O.<bucket>.C.<nuid> for chunks and $O.<bucket>.M.<name> for metadata. Clients split objects into chunks (128 KB by default in nats.go) and reassemble them on get, with a SHA-256 digest for integrity. Replication follows the stream's replica count.
Features Added in 2.11 to 2.15¶
| Release | Feature | How it works |
|---|---|---|
| 2.11 | Distributed message tracing | A Nats-Trace-Dest header makes every server on the path publish trace events (ingress, egress, mappings, account boundaries) to that subject. |
| 2.11 | Per-message TTL | Nats-TTL header expires individual messages; subject_delete_marker_ttl leaves a marker when max_age removes the last message of a subject. |
| 2.11 | Priority groups and consumer pause | Pull clients join a group; overflow and pinned_client policies decide who receives messages. pause_until stops delivery until a deadline. |
| 2.11 | Ingest rate limiting | Per-stream queues are capped (10,000 msgs / 128 MB); excess returns 429 JSStreamTooManyRequests. |
| 2.12 | Atomic batch publish | allow_atomic streams accept a batch (Nats-Batch-Id, Nats-Batch-Sequence, Nats-Batch-Commit) that commits all-or-nothing. |
| 2.12 | Counter CRDT | allow_msg_counter streams sum Nats-Incr values per subject; counters aggregate through sources. |
| 2.12 | Delayed scheduling | allow_msg_schedules streams hold a message with Nats-Schedule and publish it to Nats-Schedule-Target at the given time. |
| 2.12 | Mirror promotion and offline assets | A mirror can be promoted to primary for disaster recovery. Servers put streams that use unknown features into an offline mode after a downgrade. |
| 2.14 | Fast batch publish | allow_batched streams accept flow-controlled high-throughput batches without staging, with optional end-of-batch commit. |
| 2.14 | Recurring schedules and sampling | Schedules accept intervals or cron expressions and can sample the last message of a subject for downsampling. |
| 2.14 | Reliable WorkQueue and Interest sourcing | Sourcing uses a durable consumer with the new flow_control ack policy, so messages are acked only after the target persisted them. |
| 2.14 | Consumer reset API and feature flags | Reset delivery state to the ack floor or a sequence; feature_flags stages behaviour changes such as the v2 $JS.ACK subject format. |
| 2.15 | Desired-state meta-layer | Streams and consumers are reconciled towards a desired state, making moves and scale operations safer; moves can be cancelled ($JS.API.STREAM.CANCEL_MOVE). |
| 2.15 | Evacuate and rescue | Evacuate endpoints drain a server or a stream peer for maintenance; $JS.API.META.RESCUE temporarily lowers meta-layer quorum when nodes are permanently lost. |
| 2.15 | Backup/restore v2 | Per-message backup format with correct consumer replication; nats backup in natscli 0.5 targets it. |
Security Model¶
NATS security is decentralized and identity-first. Authentication decides which account a connection belongs to; authorization (publish and subscribe permissions) and limits come from the user and account.
Decentralized JWT Trust Chain¶
The operator signs accounts, accounts sign users, and servers trust only the operator public key. Servers never store user secrets.
flowchart TB
Op["Operator NKey (O)<br/>root of trust, kept offline"]
SK["Operator signing key"]
SysAcc["System account ($SYS)"]
AccA["Account A JWT"]
AccB["Account B JWT"]
UA1["User A1 JWT + NKey seed<br/>(.creds file)"]
UB1["User B1 JWT + NKey seed"]
Resolver["Account resolver<br/>(full / cache / memory / URL)"]
Server["nats-server<br/>trusted operator = O..."]
Op --> SK
SK -->|signs| SysAcc
SK -->|signs| AccA
SK -->|signs| AccB
AccA -->|signs| UA1
AccB -->|signs| UB1
AccA -. "nsc push" .-> Resolver
AccB -. "nsc push" .-> Resolver
Resolver --> Server
UA1 -. "CONNECT with signed nonce" .-> Server
- NKeys are Ed25519 key pairs encoded with a role prefix (
O,A,U,N,C;Smarks a seed,Xa curve key for xkey encryption). - Account JWTs carry limits (connections, subscriptions, payload, JetStream storage, leaf node connections), exports, imports and revocations.
- User JWTs carry publish and subscribe permissions, response permissions, connection types, source networks and an optional expiry.
- Resolvers deliver account JWTs to servers: the built-in
fullNATS resolver replicates JWTs between servers,cachefetches on demand,MEMORYpreloads them from config, andURLfetches over HTTP. - Auth callout (2.10) hands the CONNECT to an external service that returns a signed user JWT, which lets NATS integrate LDAP, OIDC or custom identity systems.
Authorization and Isolation¶
Permissions are subject patterns. pub.allow/deny and sub.allow/deny bound what a user can do; resp permits replies to received requests without granting a broad publish permission. Accounts are the multi-tenancy boundary: without an export and import, two accounts cannot exchange messages, even on the same server.
Encryption¶
- In transit - every listener (clients, routes, gateways, leaf nodes, MQTT, WebSocket, monitoring) has its own
tlsblock, so mTLS can be mandatory for server-to-server links and optional for clients. Since 2.12 the server uses Go's default cipher suite list and disables insecure suites unlessallow_insecure_cipher_suitesis set. - At rest - JetStream encrypts message blocks, indexes and consumer state with a per-server key (
cipher+key). Key rotation usesprev_key. On Windows the key can be sealed in the TPM (2.11).
Threat Model¶
| Threat | Mitigation |
|---|---|
| Operator key theft | Keep the operator key offline; use signing keys for routine account operations. |
| Account JWT forgery | Servers validate the signature chain up to the trusted operator key. |
| User credential theft | Short-lived user JWTs (expires), account revocations, source-network restrictions. |
| Cross-account data leakage | Accounts are isolated by design; review exports and imports, especially wildcards. |
| Slow-consumer denial of service | max_pending and write_deadline disconnect slow clients; pull consumers bound work per client. |
| JetStream storage tampering | At-rest encryption, filesystem integrity controls, monitoring $JS.EVENT.ADVISORY.>. |
| Route, gateway or leaf spoofing | mTLS on inter-server listeners. The June 2026 advisories included a route and leaf pre-auth bypass, so patch levels matter. |
| Message replay | Publish dedup with Nats-Msg-Id; application-level idempotency keys. |
| Pre-auth crashes on protocol bridges | Several 2026 advisories hit MQTT and WebSocket parsers; disable unused listeners and patch promptly. |
Advisory history and a hardening checklist are in Reference.
Performance Characteristics¶
The NATS docs publish example nats bench runs, all on one MacBook Pro M4 (10 cores, 16 GB) running nats-server 2.12.1 locally, with the client on the same machine (nats bench docs, source nats-io/nats.docs, checked 2026-09-28). The docs call them "just examples". They show relative costs, not sizing numbers for server hardware or a network.
| Workload (docs example) | Result |
|---|---|
| Core NATS publish, 1 client, 16 B, no subscriber | 14,786,683 msgs/s |
| Core NATS publish with 1 subscriber, 16 B | 4,925,767 msgs/s published, 4,928,153 msgs/s received |
| Core NATS request-reply, 1 synchronous requester | 50.87 µs average round trip |
| JetStream sync publish, memory stream, R1, 16 B | 35,734 msgs/s (27.98 µs per message) |
| JetStream batch publish (batch 1,000), memory stream, R1, 16 B | 627,430 msgs/s |
| JetStream async publish, file stream (default), R1, 128 B | 403,828 msgs/s |
| JetStream async publish, 8 clients, with 8 ordered consumers reading | 252,544 msgs/s published, 878,326 msgs/s consumed in aggregate |
The docs give no replicated (R3) or cross-cluster example. Run nats bench on your own hardware before sizing: its modes cover Core pub/sub, request-reply services, JetStream sync, async and atomic batch publishing, KV, and ordered and durable consumers, so runs can be compared under the same settings.
Governance and Licensing History¶
NATS was created by Derek Collison, first written in Ruby for Cloud Foundry and later rewritten in Go. It joined the CNCF as an incubating project on 2018-03-15, with Synadia (founded by Collison) as the main sponsor.
In April 2025 Synadia told the CNCF it wanted to take the project back, move future server releases to the Business Source License, and regain control of the trademark, domain and repositories. The CNCF refused and filed to protect the NATS trademark. On 2025-05-01 both sides announced a settlement:
- Synadia assigned its two NATS trademark registrations to the Linux Foundation.
- The nats.io domain and the
nats-ioGitHub organisation stay with the CNCF. - NATS server and clients remain Apache-2.0 under neutral CNCF governance. Synadia stays free to build commercial products on top.
NATS applied for CNCF graduation in February 2026 (cncf/toc#2042); as of 2026-09 the application is in TOC due diligence. The application cites a Trail of Bits security audit completed in May 2025.
Comparison Hooks¶
- vs Kafka - Kafka wins on the analytic streaming ecosystem (Connect, Streams, tiered storage). NATS wins on multi-tenant edge fabrics, request-reply, and operational simplicity.
- vs RabbitMQ - RabbitMQ wins on exchange routing primitives and dead-lettering. NATS wins on latency, footprint, and global topologies.
- vs Pulsar - Pulsar separates compute and storage with tiered storage. NATS keeps storage local and uses mirrors and sources for similar effects at smaller scale.
Sources¶
- NATS concepts and JetStream documentation
- nats-server source at v2.15.0 (
raft.go,gateway.go,filestore.go,stream.go) - Upgrade guides 2.11, 2.12, 2.14 and the v2.15.0 release notes
- NATS architecture and design ADRs
- CNCF and Synadia announcement (2025-05-01) and Synadia press release
- The New Stack: CNCF and Synadia reach an agreement
- NATS Messaging (Wikipedia)