Skip to content

How-to Guides

Task recipes for running RabbitMQ 4.x in production: deploying a cluster, sizing, applying best practices, securing it (TLS, OAuth 2.0, permissions), troubleshooting, upgrading to 4.3, and a Commands & Recipes section for rabbitmqctl, rabbitmq-diagnostics, rabbitmq-queues, rabbitmq-streams and rabbitmqadmin v2. Defaults and flags are listed in Reference; the reasons behind them are in Explanation.

Commands target 4.3.x

Some settings used in older guides no longer work: cluster_partition_handling has no effect since 4.3, x-queue-mode is rejected since 4.3, and the rabbitmqadmin v1 download endpoint was removed in 4.3. Recipes below use the current forms.

Deployment Patterns

Run three nodes on separate hosts or availability zones. Khepri, quorum queues (default group size 3) and streams all need a majority, so three nodes tolerate one failure. Use an odd node count; the production checklist sets a minimum of 4 CPU cores and 4 GiB RAM per node.

Five-node cluster

Five nodes tolerate two failures and spread queue leaders over more hosts. Quorum queues still default to three members, so declare with x-quorum-initial-group-size: 5 (or use CMR with target_group_size = 5) if you want five replicas; more replicas mean more replication traffic per publish.

Kubernetes (Cluster Operator)

Use the RabbitMQ Cluster Operator to declare clusters with a RabbitmqCluster resource, and the Messaging Topology Operator to manage queues, exchanges, users and policies as YAML.

apiVersion: rabbitmq.com/v1beta1
kind: RabbitmqCluster
metadata:
  name: prod
  namespace: rabbit
spec:
  replicas: 3
  resources:
    requests:
      cpu: 4
      memory: 8Gi
    limits:
      cpu: 4
      memory: 8Gi
  persistence:
    storage: 200Gi
    storageClassName: ssd
  rabbitmq:
    additionalConfig: |
      default_queue_type = quorum
      disk_free_limit.absolute = 4GB
      anonymous_login_user = none

Note

Do not add cluster_partition_handling on 4.3+; it has no effect. The operator derives the memory watermark from the container limit, so set requests equal to limits for memory.

Blue-Green migration

For upgrades that cannot be done in place (3.12 to 4.x, 3.13 with experimental Khepri, or leaving classic mirrored queues behind), build a new cluster, import definitions, and move traffic with Shovels or Federation. Since 4.2, rabbitmqadmin v2 includes commands to automate parts of a 3.13 to 4.2 Blue-Green migration (Blue-Green guide).

Sizing

Resource Guidance
CPU 4 cores minimum per node (official); quorum queues and TLS are CPU-heavy
Memory 4 GiB minimum per node (official); keep at least 30% of RAM for the OS when raising the watermark above 0.7
Disk Fast local SSD/NVMe. Quorum queues fsync their WAL; streams are I/O-bound. Avoid NAS
Free disk alarm disk_free_limit.absolute at least equal to the memory watermark (for example 4GB on an 8 GiB node)
Network Replication multiplies publish traffic by the replica count; size inter-node links accordingly
File descriptors Raise nofile well above the expected connection count (tens of thousands for busy nodes)

Best Practices

  • Default to quorum queues for durable work: default_queue_type = quorum node-wide, or --default-queue-type quorum per vhost.
  • Keep the delivery limit (default 20) and attach a dead-letter exchange via policy so poison messages are parked, not dropped.
  • Use publisher confirms and consumer acknowledgements for anything that must not be lost.
  • Set per-consumer prefetch (basic.qos); global QoS is denied by default since 4.3 and streams reject it.
  • Use streams for large fan-out or replay instead of many fanout-bound queues.
  • Federation or Shovel across WANs, not clustering; a cluster needs low-latency links.
  • One vhost per environment or tenant with tight permission regexes.
  • Enable the Prometheus plugin and scrape :15692/metrics.
  • Long-lived connections, one channel per thread; opening a connection per request causes churn.
  • Disable anonymous logins and keep guest on loopback.
  • Keep large payloads out of messages: the default max_message_size is 16 MiB since 4.0; store blobs elsewhere and send IDs.

Configure TLS

Enable TLS listeners for clients (and the same ssl_options apply to other listeners through their own prefixes, such as management.ssl.* and stream.listeners.ssl.default).

# rabbitmq.conf
listeners.tcp = none            # close plaintext AMQP
listeners.ssl.default = 5671
ssl_options.cacertfile = /etc/rabbitmq/certs/ca.pem
ssl_options.certfile   = /etc/rabbitmq/certs/server.pem
ssl_options.keyfile    = /etc/rabbitmq/certs/server.key
ssl_options.verify     = verify_peer
ssl_options.fail_if_no_peer_cert = true
ssl_options.versions.1 = tlsv1.3
ssl_options.versions.2 = tlsv1.2

Check it with rabbitmq-diagnostics listeners and openssl s_client -connect host:5671. Revocation checking (CRL/OCSP) needs separate configuration if you require it.

Encrypt inter-node traffic

Erlang distribution (25672) is plaintext by default. Point the runtime at the inet_tls distribution module (Inter-node TLS guide):

# rabbitmq-env.conf or environment
SERVER_ADDITIONAL_ERL_ARGS="-pa $ERL_SSL_PATH \
  -proto_dist inet_tls \
  -ssl_dist_opt server_certfile /etc/rabbitmq/certs/combined_keys.pem"

CLI tools need matching RABBITMQ_CTL_ERL_ARGS to reach the node.

Configure OAuth 2.0

The rabbitmq_auth_backend_oauth2 plugin validates JWTs (it does not do opaque-token introspection). Since 3.13 the simplest setup is to point it at the IdP issuer and let it discover the JWKS endpoint.

rabbitmq-plugins enable rabbitmq_auth_backend_oauth2
# rabbitmq.conf
auth_backends.1 = rabbit_auth_backend_oauth2
auth_backends.2 = rabbit_auth_backend_internal   # optional fallback
auth_oauth2.resource_server_id = rabbitmq
auth_oauth2.issuer = https://idp.example.com/realms/prod
auth_oauth2.preferred_username_claims.1 = preferred_username
# map IdP-specific scopes to RabbitMQ scopes
auth_oauth2.scope_aliases.orders-producer = rabbitmq.write:prod/orders.* rabbitmq.configure:prod/orders.*
auth_oauth2.scope_aliases.orders-consumer = rabbitmq.read:prod/orders.*

Clients send the token as the password (username may be empty). AMQP 1.0 clients can refresh the token on a live connection (4.1+). The draft 4.4.0 release notes tighten OIDC discovery validation (HTTPS endpoints, matching issuer); plan for it if your IdP is non-compliant.

Set Up Users and Permissions

rabbitmqctl add_vhost prod --default-queue-type quorum
rabbitmqctl add_user orders-svc 'use-a-generated-secret'
rabbitmqctl set_permissions -p prod orders-svc '^orders\.' '^orders\.' '^orders\.'
rabbitmqctl set_topic_permissions -p prod orders-svc amq.topic '^orders\.' '^orders\.'
rabbitmqctl set_user_tags ops-alice monitoring
rabbitmqctl delete_user guest

The permission arguments are configure, write, read (regular expressions on resource names). See Reference: Access Control for what each permission and tag covers.

Enable Auditing and Observability

  • Logs: by default in /var/log/rabbitmq/; authentication failures and permission denials are logged at info/warning.
  • Events: rabbitmq-plugins enable rabbitmq_event_exchange publishes connection, channel, user, vhost and policy events to amq.rabbitmq.event; bind a queue and ship it to a SIEM.
  • Tracing: rabbitmq_tracing (Firehose) records published and delivered messages; it is expensive, so enable it only while debugging.
  • Metrics: rabbitmq-plugins enable rabbitmq_prometheus, then scrape :15692/metrics; the official Grafana dashboards live in the rabbitmq-server repo under deps/rabbitmq_prometheus/docker/grafana. After upgrading to 4.2+, update Raft dashboards (metric names changed).
  • Secrets in parameters: Federation and Shovel URIs contain credentials; restrict who can read runtime parameters (a 2026 advisory fixed disclosure to monitoring users).

Troubleshooting

Memory alarm: publishers blocked

Symptom: publishers receive connection.blocked; the management UI shows a memory alarm.

Causes: large backlogs, many connections/channels, big management stats retention, a watermark too high for the container.

Fixes:

  • rabbitmq-diagnostics memory_breakdown to find the largest consumer of memory.
  • Drain or shorten the backlog; move heavy queues to quorum queues or streams.
  • Check the node sees the right memory limit (total_memory_available_override_value in containers).
  • Raise vm_memory_high_watermark.relative (default 0.6) only as a temporary measure.

Disk free alarm

rabbitmq-diagnostics check_alarms
df -h /var/lib/rabbitmq

Reduce stream retention (max-age, max-length-bytes via policy), remove unused queues, or grow the volume. A quorum queue stuck in a requeue loop also grows its log; check delivery-limit.

Quorum queue has no leader

Symptom: rabbitmq-queues quorum_status NAME shows no leader; publishes time out.

Cause: fewer than a majority of members reachable (for example 1 of 3).

Fix: restore the missing nodes or network. If a node is gone permanently, remove it with rabbitmqctl forget_cluster_node (4.3 removes its queue and stream members first) or rabbitmq-queues delete_member, then add a member on a healthy node. force_reset is deprecated since 4.1 and incompatible with Khepri.

Metadata operations fail during a partition

Since 4.3 (and on 4.2 clusters with Khepri), declaring queues, bindings or users fails on the minority side of a partition. This is expected. Check with rabbitmqctl cluster_status and restore connectivity; there is no partition-handling strategy to tune.

Slow consumers

rabbitmqctl list_queues name messages_ready messages_unacknowledged consumers consumer_utilisation

Low consumer utilisation with a high unacked count means consumers are the bottleneck: add consumers, lower per-message work, or tune prefetch. On 4.3 quorum queues, consumer-timeout returns messages held too long.

Node.js clients cannot connect after upgrading to 4.1+

amqplib older than 0.10.7 uses a 4096-byte pre-auth frame_max; 4.1 requires at least 8192. Upgrade the library.

MQTT or Web MQTT clients disconnecting

  • Confirm rabbitmq_mqtt / rabbitmq_web_mqtt are enabled on every node and listeners are configured (mqtt.listeners.tcp.default, mqtt.listeners.ssl.default).
  • Since 4.1 the default MQTT max packet size is 16 MiB; larger packets are rejected.
  • Since 4.0 mqtt.default_user is gone; use anonymous_login_user or real credentials.

Upgrade to 4.3

flowchart TD
    Now{"Current series?"} -->|"3.12.x or older"| BG["Blue-Green migration<br/>to a new 4.2+ cluster"]
    Now -->|"3.13.x with experimental Khepri"| BG
    Now -->|"3.13.x (Mnesia)"| To42["Enable all feature flags,<br/>rolling upgrade to latest 4.2.x"]
    Now -->|"4.0.x or 4.1.x"| To42
    Now -->|"4.2.x"| Khepri["rabbitmqctl enable_feature_flag all<br/>(includes khepri_db)"]
    To42 --> Khepri
    Khepri --> Erl{"Erlang 27+ on every node?"}
    Erl -->|no| UpErl["Upgrade Erlang first"]
    UpErl --> Roll
    Erl -->|yes| Roll["Rolling upgrade node by node<br/>to latest 4.3.x"]
    Roll --> Clean["Remove cluster_partition_handling keys,<br/>check deprecated features,<br/>rabbitmq-queues rebalance all"]
  1. Read the 4.3.0 release notes. Only 4.2.x upgrades in place to 4.3.x.
  2. Check applications for features now denied by default: transient non-exclusive queues, global QoS, x-queue-mode, AMQP 1.0 v1 addresses. Opt back in temporarily with deprecated_features.permit.<name> = true on all nodes before upgrading.
  3. Enable every stable feature flag: rabbitmqctl enable_feature_flag all, then confirm khepri_db is enabled with rabbitmqctl list_feature_flags.
  4. For each node: rabbitmq-diagnostics check_if_node_is_quorum_critical (must pass), rabbitmq-upgrade drain, stop, upgrade the package, start, rabbitmq-upgrade revive, then rabbitmq-diagnostics check_running before moving on. Automation can use rabbitmq-upgrade await_online_quorum_plus_one.
  5. Keep the mixed-version window to hours, not days.
  6. Afterwards run rabbitmq-queues rebalance all and remove partition-handling keys from rabbitmq.conf.

Migrating off classic mirrored queues

After an upgrade to 4.x, mirroring policies silently stop replicating. Declare replacement quorum queues (queue type is immutable), move consumers, and use a Shovel to drain the old classic queues. See Migrate mirrored classic queues to quorum queues.

Cost Analysis

Cost Driver
Compute Erlang VM baseline is moderate; quorum queues and TLS add CPU per message
Storage Quorum queue WAL fsyncs and stream segments; provision SSD/NVMe and headroom above disk_free_limit
Memory Connections, channels and quorum queue in-memory index; 4.3 roughly halves per-message overhead in many cases
Network egress Replication traffic within the cluster; Federation/Shovel duplicate traffic across regions
Tanzu RabbitMQ Commercial license from Broadcom; required for patches to out-of-community-support series (pricing not public)
Managed services CloudAMQP (per plan), Amazon MQ for RabbitMQ (per broker-hour and storage)

Commands & Recipes

Bootstrap and cluster

# On node 1
rabbitmq-diagnostics status

# On node 2: join node 1's cluster
rabbitmqctl stop_app
rabbitmqctl reset
rabbitmqctl join_cluster rabbit@node1
rabbitmqctl start_app
rabbitmqctl cluster_status

Declare a quorum queue, exchange and binding (rabbitmqadmin v2)

rabbitmqadmin --vhost prod queues declare --name orders --type quorum --durable true
rabbitmqadmin --vhost prod exchanges declare --name orders.x --type topic --durable true
rabbitmqadmin --vhost prod bindings declare --source orders.x --destination-type queue \
  --destination orders --routing-key 'orders.created.*'
# delivery limit, length cap and DLX via policy (preferred over queue arguments)
rabbitmqadmin --vhost prod policies declare --name orders-limits --pattern '^orders$' \
  --definition '{"delivery-limit":5,"max-length":1000000,"dead-letter-exchange":"orders.dlx"}' \
  --priority 1 --apply-to quorum_queues

Declare a stream and a super stream

rabbitmqadmin --vhost prod streams declare --name events --expiration 7D
rabbitmqctl set_policy -p prod events-retention '^events$' \
  '{"max-length-bytes":50000000000}' --apply-to streams
rabbitmq-streams add_super_stream invoices --partitions 3 --vhost prod

Federation upstream

rabbitmqctl set_parameter -p prod federation-upstream us-prod \
  '{"uri":"amqps://fed-user:secret@us-prod.example.com:5671"}'
rabbitmqctl set_policy -p prod federate-orders '^orders\.' \
  '{"federation-upstream-set":"all"}' --apply-to exchanges

Diagnostics

rabbitmq-diagnostics status
rabbitmq-diagnostics memory_breakdown
rabbitmq-diagnostics check_alarms
rabbitmq-diagnostics check_running
rabbitmq-diagnostics observer            # text-mode observer (top-like)
rabbitmqctl list_queues name type messages messages_ready consumers
rabbitmqctl list_connections user host channels state
rabbitmq-queues quorum_status orders
rabbitmq-streams stream_status events
rabbitmqadmin health_check local_alarms

Prometheus

rabbitmq-plugins enable rabbitmq_prometheus
curl -s http://node:15692/metrics | head

perf-test

docker run -it --rm pivotalrabbitmq/perf-test:latest \
  --uri amqp://app:pwd@host:5672 \
  --producers 10 --consumers 10 \
  --rate 10000 --confirm 100 \
  --queue-pattern 'q-%d' --queue-pattern-from 1 --queue-pattern-to 50 \
  --quorum-queue

Cross-references

  • Explanation: queue types, Khepri and the security model behind these recipes.
  • Reference: config defaults, ports, upgrade paths, hardening checklist.
  • Messaging: comparisons with Kafka, NATS, Redpanda and Pulsar.

Sources