Skip to content

How-to Guides

Task-oriented recipes for running NATS Server 2.15 and JetStream: choosing a topology, configuring clusters, gateways and leaf nodes, securing listeners, creating streams, KV buckets and object stores, running on Kubernetes, monitoring, troubleshooting, and upgrading. Defaults and config keys are listed in Reference; the reasons behind them are in Explanation.

Commands target current tooling

Commands use nats-server 2.15, nats CLI 0.5 and nsc 2.15. natscli 0.5 moved stream backups to nats backup, and nats bench uses subcommands (pub, sub, js ...).

Choose a Deployment Topology

Need Topology Notes
One site, HA JetStream Single cluster of 3 (or 5) servers R3 streams tolerate one server loss; R5 tolerates two.
Several regions, one subject space Supercluster (clusters joined by gateways) Place streams per cluster with placement; use mirrors for cross-region reads.
Edge sites, customer premises, firewalled networks Hub cluster + leaf nodes Leaf dials out on 7422; give each site its own JetStream domain to keep working offline.
Kubernetes Official Helm chart (StatefulSet) One server per pod, persistent volume per pod, never autoscale JetStream members.

The default single-cluster footprint is three servers in a full mesh of routes, with clients spread across all of them.

flowchart LR
    subgraph C1["Cluster C1 (JetStream enabled)"]
        n1["nats-server n1"]
        n2["nats-server n2"]
        n3["nats-server n3"]
        n1 <-->|"route :6222"| n2
        n2 <-->|"route :6222"| n3
        n1 <-->|"route :6222"| n3
    end
    Svc["Services (nats.go, async-nats, jnats)"] -->|":4222"| n1
    Svc --> n2
    Svc --> n3
    Leaf["Leaf node (edge site)"] -->|"leaf :7422"| n2

Configure a JetStream Cluster

Save as /etc/nats/nats.conf on each server, changing server_name:

server_name: n1
port: 4222
http_port: 8222

jetstream {
  store_dir: /data/jetstream
  max_memory_store: 4GB
  max_file_store: 500GB
}

cluster {
  name: C1
  port: 6222
  routes: [
    nats-route://n1.example.internal:6222
    nats-route://n2.example.internal:6222
    nats-route://n3.example.internal:6222
  ]
  connect_backoff: true   # 2.12+: exponential reconnect backoff
}

For a throwaway local test of the same shape (no auth, dev only):

nats-server -js -sd /tmp/n1 -p 4222 --server_name n1 --cluster_name C1 --cluster nats://127.0.0.1:6222 --routes nats://127.0.0.1:6222,nats://127.0.0.1:6223,nats://127.0.0.1:6224 &
nats-server -js -sd /tmp/n2 -p 4223 --server_name n2 --cluster_name C1 --cluster nats://127.0.0.1:6223 --routes nats://127.0.0.1:6222,nats://127.0.0.1:6223,nats://127.0.0.1:6224 &
nats-server -js -sd /tmp/n3 -p 4224 --server_name n3 --cluster_name C1 --cluster nats://127.0.0.1:6224 --routes nats://127.0.0.1:6222,nats://127.0.0.1:6223,nats://127.0.0.1:6224 &
nats account info --server nats://127.0.0.1:4222   # confirms JetStream is enabled for the account

Validate a config file and print its digest (2.11+) before rolling it out:

nats-server -c /etc/nats/nats.conf -t

Connect Clusters with Gateways

Add a gateway block to every server of each cluster. The gateway name must equal the cluster name:

gateway {
  name: C1
  port: 7222
  gateways: [
    { name: C1, urls: ["nats://c1-n1.example.internal:7222"] }
    { name: C2, urls: ["nats://c2-n1.example.internal:7222"] }
  ]
}

Then pin streams to a region, for example nats stream add ORDERS_EU --cluster C2 ..., and mirror them elsewhere if other regions need local reads.

Attach a Leaf Node

On the edge server, dial out to the hub with the credentials of the hub account the site should join:

leafnodes {
  remotes: [
    {
      url: "tls://hub.example.internal:7422"
      credentials: "/etc/nats/site-17.creds"
    }
  ]
}

jetstream {
  store_dir: /data/jetstream
  domain: site17        # local JetStream keeps working when the link is down
}

On the hub, enable the listener with leafnodes { port: 7422 } and, for large fleets whose sites never talk to each other, isolate_leafnode_interest: true (2.12+). Since 2.14 remotes can be added or removed with nats-server --signal reload.

Secure the Deployment

Bootstrap Decentralized Auth with nsc

nsc add operator --generate-signing-key --sys --name DEMO
nsc edit operator --require-signing-keys --account-jwt-server-url nats://n1.example.internal:4222
nsc add account APP
nsc edit account APP --sk generate --js-mem-storage 1G --js-disk-storage 50G
nsc add user --account APP service
nsc generate creds --account APP --name service > app-service.creds
nsc generate config --nats-resolver --sys-account SYS > resolver.conf

Include resolver.conf in the server config, start the servers, then push account JWTs:

nsc push --all

When permissions change with nsc edit, run nsc push --account APP again; servers using the full resolver pick up the new JWT without a restart.

Enable TLS and mTLS

Each listener has its own tls block. Require client certificates on inter-server links at minimum:

tls {
  cert_file: "/etc/nats/certs/server.pem"
  key_file:  "/etc/nats/certs/server-key.pem"
  ca_file:   "/etc/nats/certs/ca.pem"
  verify:    true
}

cluster {
  tls {
    cert_file: "/etc/nats/certs/route.pem"
    key_file:  "/etc/nats/certs/route-key.pem"
    ca_file:   "/etc/nats/certs/ca.pem"
    verify:    true
  }
}

Use verify_and_map: true instead of verify to map the certificate subject to a user. Since 2.12 insecure cipher suites are off by default; cipher_suites only affects TLS 1.2 in Go, since TLS 1.3 suites are not configurable.

Encrypt JetStream at Rest

jetstream {
  store_dir: /data/jetstream
  cipher: chachapoly          # or aes
  key: $JS_KEY                # from an environment variable, never inline in Git
}

To rotate, set the new value as key and the old one as prev_key, then restart; data is re-encrypted as it is rewritten.

Manage Streams, KV and Object Store

# Replicated stream over orders.> with dedup window
nats stream add ORDERS \
  --subjects "orders.>" --storage file --replicas 3 \
  --retention limits --max-age 720h --max-bytes 50GB \
  --discard old --dupe-window 2m --defaults

# Durable pull consumer
nats consumer add ORDERS workers \
  --pull --filter "orders.created.>" \
  --ack explicit --max-deliver 5 --replay instant --defaults

# KV bucket with history and TTL
nats kv add SESSIONS --replicas 3 --history 5 --ttl 24h
nats kv put SESSIONS user.42 '{"role":"admin"}'
nats kv get SESSIONS user.42
nats kv watch SESSIONS 'user.>'

# Object store
nats object add FIRMWARE --replicas 3
nats object put FIRMWARE ./build/firmware-v1.bin --name firmware-v1.bin
nats object get FIRMWARE firmware-v1.bin
nats object info FIRMWARE

Guidelines:

  • Retention - pick one of max_age, max_bytes, max_msgs as the dominant limit per stream and keep the others as safety nets.
  • Replicas - R1 for development and edge, R3 for production, R5 only for critical metadata or KV.
  • Consumers - prefer pull consumers for worker pools. From 2.15 a stream allows 1,000 consumers by default (jetstream { limits { default_max_consumers } } or the stream's max_consumers to change).
  • Accounts - one account per trust boundary (tenant, environment, team), with exports and imports for the few legitimate crossings.

Size and Tune a Cluster

Starting-point guidance (not vendor-published figures); validate with nats bench on your hardware:

Resource Guidance
CPU 4-8 vCPUs per server is typical. Core NATS is rarely CPU-bound; JetStream R3 and TLS add load.
Memory 8 GB baseline for JetStream servers. Raise for large interest graphs and stream caches; set GOMEMLIMIT below the container limit (2.12+).
Disk NVMe for store_dir. Watch fsync and write latency.
Network 10 GbE or more for high-throughput replicated streams. Gateways tolerate WAN latency.
Server count Odd numbers (3, 5) for Raft quorum.

Tuning checklist (keys and defaults in Reference):

  • Subject design - stable token order such as region.tenant.entity.action.id; no PII in subject tokens, because subjects appear in logs, traces and permissions.
  • max_payload - keep the 1 MB default; use Object Store for large payloads.
  • write_deadline / max_pending - lower the deadline to drop slow consumers faster; raise max_pending only with memory headroom.
  • cluster { pool_size } - raise from 3 for heavy cross-route fan-out, or pin busy accounts to dedicated routes.
  • cluster { no_advertise: true } - hide internal addresses from client discovery when clients connect through a load balancer.
  • jetstream { max_memory_store, max_file_store } - cap per-server JetStream resources so one stream cannot exhaust the node.
  • Account limits - set JetStream storage limits per account (nsc edit account --js-disk-storage).
  • Monitoring user - create a dedicated system-account user for monitoring tools; never share it with applications.
  • Consistent JetStream membership - avoid clusters where only some servers run JetStream unless you place assets deliberately with tags.

Back Up and Restore Streams

natscli 0.5 moved backups to the nats backup command; nats stream backup is hidden and deprecated.

nats backup stream ORDERS /backups/ORDERS
nats backup info /backups/ORDERS           # backups taken from 2.15+ servers
nats backup validate /backups/ORDERS
nats backup restore stream /backups/ORDERS
nats backup account /backups/all-streams   # every stream in the account

Back up the nsc store (~/.local/share/nats/nsc by default) and operator keys separately, offline.

Run on Kubernetes

Install the server with the official Helm chart (chart 2.15.0 ships nats:2.15.0-alpine):

helm repo add nats https://nats-io.github.io/k8s/helm/charts/
helm repo update
helm upgrade --install nats nats/nats \
  --set config.cluster.enabled=true \
  --set config.cluster.replicas=3 \
  --set config.jetstream.enabled=true \
  --set config.jetstream.fileStore.pvc.size=200Gi \
  --set config.jetstream.fileStore.pvc.storageClassName=ssd \
  -n nats --create-namespace --wait

Manage JetStream assets declaratively with NACK. KeyValue, ObjectStore and Account resources require the --control-loop controller:

helm upgrade --install nack nats/nack -n nats \
  --set jetstream.nats.url=nats://nats.nats.svc.cluster.local:4222 \
  --set jetstream.additionalArgs={--control-loop} --wait
kubectl apply -f https://raw.githubusercontent.com/nats-io/nack/main/deploy/examples/stream.yml
kubectl apply -f https://raw.githubusercontent.com/nats-io/nack/main/deploy/examples/consumer_pull.yml

Resources use apiVersion: jetstream.nats.io/v1beta2 and are exclusively owned by NACK: changes made with the nats CLI are reverted.

Monitor NATS

nats-server exposes JSON endpoints on the monitoring port (/varz, /connz, /routez, /gatewayz, /leafz, /jsz, /healthz). Convert them to Prometheus metrics with the exporter (the Helm chart can run it as a sidecar):

prometheus-nats-exporter -varz -connz -routez -gatewayz -leafz -healthz -jsz=all http://nats-1:8222
# scrape http://<exporter>:7777/metrics

Useful live views:

nats-top -s nats://nats-1:4222
nats server report jetstream       # system account
nats stream report
nats consumer report ORDERS
curl -s 'http://nats-1:8222/healthz?js-enabled-only=true'

Use healthz for Kubernetes probes. Since 2.11 js-server-only no longer checks the meta leader; use js-meta-only if you want that check.

Benchmark

nats bench pub test --msgs 10000000 --clients 2 --no-progress
nats bench sub test --msgs 10000000 --clients 2 --no-progress
nats bench js pub async bench.js --msgs 1000000 --size 512 --create --replicas 3 --storage file

Always benchmark on production-like disks and networks; results on laptops say little about R3 file streams.

Troubleshoot

Slow Consumers

Symptom: logs show Slow Consumer Detected and the client is disconnected; gnatsd_varz_slow_consumers rises.

Fixes:

  • Move blocking work (database calls, HTTP) out of message callbacks.
  • Switch high-volume work to JetStream pull consumers with bounded Fetch batches.
  • Check nats-top for pending bytes; raise max_pending only with care.

Stream Replica Lagging

Symptom: nats stream info ORDERS shows a replica as current: false or with a growing lag.

Fixes:

  • Check disk latency (iostat -x 1) and free space in store_dir.
  • Move leadership away from a slow server: nats stream cluster step-down ORDERS.
  • Rebalance leaders across the cluster: nats stream cluster balance.
  • For planned maintenance on 2.15, evacuate the server first: nats server cluster evacuate <server>.

Stream Frozen after I/O Errors (2.14+)

Symptom: a stream stops accepting writes, the health check fails, and logs mention write error.

Fix: fix the underlying disk problem, then restart the server; replicated streams keep serving from other replicas in the meantime.

JetStream Request Rejected (2.12+)

Symptom: clients get errors and logs show Invalid JetStream request ... json: unknown field.

Fix: upgrade the client library. As a temporary workaround, set jetstream { strict: false }.

429 Too Many Requests on Publish (2.11+)

Symptom: JSStreamTooManyRequests or IPQ len limit reached warnings.

Fix: publish with PubAck (async publish) rather than Core NATS fire-and-forget into streams, or raise max_buffered_msgs / max_buffered_size.

Account Permissions Not Applied

Symptom: after nsc edit, the server still enforces old permissions.

Fix: nsc push --account APP, then confirm with nats account info using the user's credentials.

Apparent Split Brain

JetStream uses Raft, so two leaders for one stream cannot commit concurrently. Apparent split brain is usually two leaf domains or two clusters using the same stream names. Check with nats server report jetstream from a system-account user. For permanent loss of a majority of meta-layer peers on 2.15, see nats server cluster rescue (Reference, $JS.API.META.RESCUE).

Upgrade Safely

  1. Read the upgrade guide for every minor line you cross (2.11, 2.12, 2.14, 2.15) and the release notes of the target patch.
  2. Upgrade one minor line at a time. For 2.15, every peer must already run 2.14.0 or later, because 2.15 enables js_raft_delete_range and older peers panic on the new Raft entry.
  3. Roll servers one by one. Put each into lame duck mode so clients reconnect elsewhere before shutdown:
nats-server --signal ldm=/var/run/nats/nats.pid
  1. Wait until healthz is green and nats server report jetstream shows all replicas current before moving to the next server.
  2. If you use account exports or permissions on $JS.ACK.<stream>.> or $JS.FC.<stream>.>, broaden them to cover the v2 subject format, then test with feature_flags { js_ack_fc_v2: true }. The switch was announced for 2.15 but is still off by default in 2.15.0.
  3. Before downgrading below 2.14 remove feature_flags from the config; from 2.12 only downgrade to 2.11.9 or later.

Cost Considerations

Cost Driver
Compute Small: a three-server cluster fits on modest VMs for most workloads.
Storage JetStream dominates. Mirror old data to a cheaper cluster if replay speed is not needed.
Network egress Gateway and leaf links carry cross-region and edge traffic.
Managed service Synadia Cloud and Synadia Platform are priced by Synadia; see the vendor site.

Sources