How-to Guides¶
Task-oriented recipes for running NATS Server 2.15 and JetStream: choosing a topology, configuring clusters, gateways and leaf nodes, securing listeners, creating streams, KV buckets and object stores, running on Kubernetes, monitoring, troubleshooting, and upgrading. Defaults and config keys are listed in Reference; the reasons behind them are in Explanation.
Commands target current tooling
Commands use nats-server 2.15, nats CLI 0.5 and nsc 2.15. natscli 0.5 moved stream backups to nats backup, and nats bench uses subcommands (pub, sub, js ...).
Choose a Deployment Topology¶
| Need | Topology | Notes |
|---|---|---|
| One site, HA JetStream | Single cluster of 3 (or 5) servers | R3 streams tolerate one server loss; R5 tolerates two. |
| Several regions, one subject space | Supercluster (clusters joined by gateways) | Place streams per cluster with placement; use mirrors for cross-region reads. |
| Edge sites, customer premises, firewalled networks | Hub cluster + leaf nodes | Leaf dials out on 7422; give each site its own JetStream domain to keep working offline. |
| Kubernetes | Official Helm chart (StatefulSet) | One server per pod, persistent volume per pod, never autoscale JetStream members. |
The default single-cluster footprint is three servers in a full mesh of routes, with clients spread across all of them.
flowchart LR
subgraph C1["Cluster C1 (JetStream enabled)"]
n1["nats-server n1"]
n2["nats-server n2"]
n3["nats-server n3"]
n1 <-->|"route :6222"| n2
n2 <-->|"route :6222"| n3
n1 <-->|"route :6222"| n3
end
Svc["Services (nats.go, async-nats, jnats)"] -->|":4222"| n1
Svc --> n2
Svc --> n3
Leaf["Leaf node (edge site)"] -->|"leaf :7422"| n2
Configure a JetStream Cluster¶
Save as /etc/nats/nats.conf on each server, changing server_name:
server_name: n1
port: 4222
http_port: 8222
jetstream {
store_dir: /data/jetstream
max_memory_store: 4GB
max_file_store: 500GB
}
cluster {
name: C1
port: 6222
routes: [
nats-route://n1.example.internal:6222
nats-route://n2.example.internal:6222
nats-route://n3.example.internal:6222
]
connect_backoff: true # 2.12+: exponential reconnect backoff
}
For a throwaway local test of the same shape (no auth, dev only):
nats-server -js -sd /tmp/n1 -p 4222 --server_name n1 --cluster_name C1 --cluster nats://127.0.0.1:6222 --routes nats://127.0.0.1:6222,nats://127.0.0.1:6223,nats://127.0.0.1:6224 &
nats-server -js -sd /tmp/n2 -p 4223 --server_name n2 --cluster_name C1 --cluster nats://127.0.0.1:6223 --routes nats://127.0.0.1:6222,nats://127.0.0.1:6223,nats://127.0.0.1:6224 &
nats-server -js -sd /tmp/n3 -p 4224 --server_name n3 --cluster_name C1 --cluster nats://127.0.0.1:6224 --routes nats://127.0.0.1:6222,nats://127.0.0.1:6223,nats://127.0.0.1:6224 &
nats account info --server nats://127.0.0.1:4222 # confirms JetStream is enabled for the account
Validate a config file and print its digest (2.11+) before rolling it out:
Connect Clusters with Gateways¶
Add a gateway block to every server of each cluster. The gateway name must equal the cluster name:
gateway {
name: C1
port: 7222
gateways: [
{ name: C1, urls: ["nats://c1-n1.example.internal:7222"] }
{ name: C2, urls: ["nats://c2-n1.example.internal:7222"] }
]
}
Then pin streams to a region, for example nats stream add ORDERS_EU --cluster C2 ..., and mirror them elsewhere if other regions need local reads.
Attach a Leaf Node¶
On the edge server, dial out to the hub with the credentials of the hub account the site should join:
leafnodes {
remotes: [
{
url: "tls://hub.example.internal:7422"
credentials: "/etc/nats/site-17.creds"
}
]
}
jetstream {
store_dir: /data/jetstream
domain: site17 # local JetStream keeps working when the link is down
}
On the hub, enable the listener with leafnodes { port: 7422 } and, for large fleets whose sites never talk to each other, isolate_leafnode_interest: true (2.12+). Since 2.14 remotes can be added or removed with nats-server --signal reload.
Secure the Deployment¶
Bootstrap Decentralized Auth with nsc¶
nsc add operator --generate-signing-key --sys --name DEMO
nsc edit operator --require-signing-keys --account-jwt-server-url nats://n1.example.internal:4222
nsc add account APP
nsc edit account APP --sk generate --js-mem-storage 1G --js-disk-storage 50G
nsc add user --account APP service
nsc generate creds --account APP --name service > app-service.creds
nsc generate config --nats-resolver --sys-account SYS > resolver.conf
Include resolver.conf in the server config, start the servers, then push account JWTs:
When permissions change with nsc edit, run nsc push --account APP again; servers using the full resolver pick up the new JWT without a restart.
Enable TLS and mTLS¶
Each listener has its own tls block. Require client certificates on inter-server links at minimum:
tls {
cert_file: "/etc/nats/certs/server.pem"
key_file: "/etc/nats/certs/server-key.pem"
ca_file: "/etc/nats/certs/ca.pem"
verify: true
}
cluster {
tls {
cert_file: "/etc/nats/certs/route.pem"
key_file: "/etc/nats/certs/route-key.pem"
ca_file: "/etc/nats/certs/ca.pem"
verify: true
}
}
Use verify_and_map: true instead of verify to map the certificate subject to a user. Since 2.12 insecure cipher suites are off by default; cipher_suites only affects TLS 1.2 in Go, since TLS 1.3 suites are not configurable.
Encrypt JetStream at Rest¶
jetstream {
store_dir: /data/jetstream
cipher: chachapoly # or aes
key: $JS_KEY # from an environment variable, never inline in Git
}
To rotate, set the new value as key and the old one as prev_key, then restart; data is re-encrypted as it is rewritten.
Manage Streams, KV and Object Store¶
# Replicated stream over orders.> with dedup window
nats stream add ORDERS \
--subjects "orders.>" --storage file --replicas 3 \
--retention limits --max-age 720h --max-bytes 50GB \
--discard old --dupe-window 2m --defaults
# Durable pull consumer
nats consumer add ORDERS workers \
--pull --filter "orders.created.>" \
--ack explicit --max-deliver 5 --replay instant --defaults
# KV bucket with history and TTL
nats kv add SESSIONS --replicas 3 --history 5 --ttl 24h
nats kv put SESSIONS user.42 '{"role":"admin"}'
nats kv get SESSIONS user.42
nats kv watch SESSIONS 'user.>'
# Object store
nats object add FIRMWARE --replicas 3
nats object put FIRMWARE ./build/firmware-v1.bin --name firmware-v1.bin
nats object get FIRMWARE firmware-v1.bin
nats object info FIRMWARE
Guidelines:
- Retention - pick one of
max_age,max_bytes,max_msgsas the dominant limit per stream and keep the others as safety nets. - Replicas - R1 for development and edge, R3 for production, R5 only for critical metadata or KV.
- Consumers - prefer pull consumers for worker pools. From 2.15 a stream allows 1,000 consumers by default (
jetstream { limits { default_max_consumers } }or the stream'smax_consumersto change). - Accounts - one account per trust boundary (tenant, environment, team), with exports and imports for the few legitimate crossings.
Size and Tune a Cluster¶
Starting-point guidance (not vendor-published figures); validate with nats bench on your hardware:
| Resource | Guidance |
|---|---|
| CPU | 4-8 vCPUs per server is typical. Core NATS is rarely CPU-bound; JetStream R3 and TLS add load. |
| Memory | 8 GB baseline for JetStream servers. Raise for large interest graphs and stream caches; set GOMEMLIMIT below the container limit (2.12+). |
| Disk | NVMe for store_dir. Watch fsync and write latency. |
| Network | 10 GbE or more for high-throughput replicated streams. Gateways tolerate WAN latency. |
| Server count | Odd numbers (3, 5) for Raft quorum. |
Tuning checklist (keys and defaults in Reference):
- Subject design - stable token order such as
region.tenant.entity.action.id; no PII in subject tokens, because subjects appear in logs, traces and permissions. max_payload- keep the 1 MB default; use Object Store for large payloads.write_deadline/max_pending- lower the deadline to drop slow consumers faster; raisemax_pendingonly with memory headroom.cluster { pool_size }- raise from 3 for heavy cross-route fan-out, or pin busy accounts to dedicated routes.cluster { no_advertise: true }- hide internal addresses from client discovery when clients connect through a load balancer.jetstream { max_memory_store, max_file_store }- cap per-server JetStream resources so one stream cannot exhaust the node.- Account limits - set JetStream storage limits per account (
nsc edit account --js-disk-storage). - Monitoring user - create a dedicated system-account user for monitoring tools; never share it with applications.
- Consistent JetStream membership - avoid clusters where only some servers run JetStream unless you place assets deliberately with tags.
Back Up and Restore Streams¶
natscli 0.5 moved backups to the nats backup command; nats stream backup is hidden and deprecated.
nats backup stream ORDERS /backups/ORDERS
nats backup info /backups/ORDERS # backups taken from 2.15+ servers
nats backup validate /backups/ORDERS
nats backup restore stream /backups/ORDERS
nats backup account /backups/all-streams # every stream in the account
Back up the nsc store (~/.local/share/nats/nsc by default) and operator keys separately, offline.
Run on Kubernetes¶
Install the server with the official Helm chart (chart 2.15.0 ships nats:2.15.0-alpine):
helm repo add nats https://nats-io.github.io/k8s/helm/charts/
helm repo update
helm upgrade --install nats nats/nats \
--set config.cluster.enabled=true \
--set config.cluster.replicas=3 \
--set config.jetstream.enabled=true \
--set config.jetstream.fileStore.pvc.size=200Gi \
--set config.jetstream.fileStore.pvc.storageClassName=ssd \
-n nats --create-namespace --wait
Manage JetStream assets declaratively with NACK. KeyValue, ObjectStore and Account resources require the --control-loop controller:
helm upgrade --install nack nats/nack -n nats \
--set jetstream.nats.url=nats://nats.nats.svc.cluster.local:4222 \
--set jetstream.additionalArgs={--control-loop} --wait
kubectl apply -f https://raw.githubusercontent.com/nats-io/nack/main/deploy/examples/stream.yml
kubectl apply -f https://raw.githubusercontent.com/nats-io/nack/main/deploy/examples/consumer_pull.yml
Resources use apiVersion: jetstream.nats.io/v1beta2 and are exclusively owned by NACK: changes made with the nats CLI are reverted.
Monitor NATS¶
nats-server exposes JSON endpoints on the monitoring port (/varz, /connz, /routez, /gatewayz, /leafz, /jsz, /healthz). Convert them to Prometheus metrics with the exporter (the Helm chart can run it as a sidecar):
prometheus-nats-exporter -varz -connz -routez -gatewayz -leafz -healthz -jsz=all http://nats-1:8222
# scrape http://<exporter>:7777/metrics
Useful live views:
nats-top -s nats://nats-1:4222
nats server report jetstream # system account
nats stream report
nats consumer report ORDERS
curl -s 'http://nats-1:8222/healthz?js-enabled-only=true'
Use healthz for Kubernetes probes. Since 2.11 js-server-only no longer checks the meta leader; use js-meta-only if you want that check.
Benchmark¶
nats bench pub test --msgs 10000000 --clients 2 --no-progress
nats bench sub test --msgs 10000000 --clients 2 --no-progress
nats bench js pub async bench.js --msgs 1000000 --size 512 --create --replicas 3 --storage file
Always benchmark on production-like disks and networks; results on laptops say little about R3 file streams.
Troubleshoot¶
Slow Consumers¶
Symptom: logs show Slow Consumer Detected and the client is disconnected; gnatsd_varz_slow_consumers rises.
Fixes:
- Move blocking work (database calls, HTTP) out of message callbacks.
- Switch high-volume work to JetStream pull consumers with bounded
Fetchbatches. - Check
nats-topfor pending bytes; raisemax_pendingonly with care.
Stream Replica Lagging¶
Symptom: nats stream info ORDERS shows a replica as current: false or with a growing lag.
Fixes:
- Check disk latency (
iostat -x 1) and free space instore_dir. - Move leadership away from a slow server:
nats stream cluster step-down ORDERS. - Rebalance leaders across the cluster:
nats stream cluster balance. - For planned maintenance on 2.15, evacuate the server first:
nats server cluster evacuate <server>.
Stream Frozen after I/O Errors (2.14+)¶
Symptom: a stream stops accepting writes, the health check fails, and logs mention write error.
Fix: fix the underlying disk problem, then restart the server; replicated streams keep serving from other replicas in the meantime.
JetStream Request Rejected (2.12+)¶
Symptom: clients get errors and logs show Invalid JetStream request ... json: unknown field.
Fix: upgrade the client library. As a temporary workaround, set jetstream { strict: false }.
429 Too Many Requests on Publish (2.11+)¶
Symptom: JSStreamTooManyRequests or IPQ len limit reached warnings.
Fix: publish with PubAck (async publish) rather than Core NATS fire-and-forget into streams, or raise max_buffered_msgs / max_buffered_size.
Account Permissions Not Applied¶
Symptom: after nsc edit, the server still enforces old permissions.
Fix: nsc push --account APP, then confirm with nats account info using the user's credentials.
Apparent Split Brain¶
JetStream uses Raft, so two leaders for one stream cannot commit concurrently. Apparent split brain is usually two leaf domains or two clusters using the same stream names. Check with nats server report jetstream from a system-account user. For permanent loss of a majority of meta-layer peers on 2.15, see nats server cluster rescue (Reference, $JS.API.META.RESCUE).
Upgrade Safely¶
- Read the upgrade guide for every minor line you cross (2.11, 2.12, 2.14, 2.15) and the release notes of the target patch.
- Upgrade one minor line at a time. For 2.15, every peer must already run 2.14.0 or later, because 2.15 enables
js_raft_delete_rangeand older peers panic on the new Raft entry. - Roll servers one by one. Put each into lame duck mode so clients reconnect elsewhere before shutdown:
- Wait until
healthzis green andnats server report jetstreamshows all replicas current before moving to the next server. - If you use account exports or permissions on
$JS.ACK.<stream>.>or$JS.FC.<stream>.>, broaden them to cover the v2 subject format, then test withfeature_flags { js_ack_fc_v2: true }. The switch was announced for 2.15 but is still off by default in 2.15.0. - Before downgrading below 2.14 remove
feature_flagsfrom the config; from 2.12 only downgrade to 2.11.9 or later.
Cost Considerations¶
| Cost | Driver |
|---|---|
| Compute | Small: a three-server cluster fits on modest VMs for most workloads. |
| Storage | JetStream dominates. Mirror old data to a cheaper cluster if replay speed is not needed. |
| Network egress | Gateway and leaf links carry cross-region and edge traffic. |
| Managed service | Synadia Cloud and Synadia Platform are priced by Synadia; see the vendor site. |