How-to Guides¶
Task recipes for running RabbitMQ 4.x in production: deploying a cluster, sizing, applying best practices, securing it (TLS, OAuth 2.0, permissions), troubleshooting, upgrading to 4.3, and a Commands & Recipes section for rabbitmqctl, rabbitmq-diagnostics, rabbitmq-queues, rabbitmq-streams and rabbitmqadmin v2. Defaults and flags are listed in Reference; the reasons behind them are in Explanation.
Commands target 4.3.x
Some settings used in older guides no longer work: cluster_partition_handling has no effect since 4.3, x-queue-mode is rejected since 4.3, and the rabbitmqadmin v1 download endpoint was removed in 4.3. Recipes below use the current forms.
Deployment Patterns¶
Three-node cluster (recommended baseline)¶
Run three nodes on separate hosts or availability zones. Khepri, quorum queues (default group size 3) and streams all need a majority, so three nodes tolerate one failure. Use an odd node count; the production checklist sets a minimum of 4 CPU cores and 4 GiB RAM per node.
Five-node cluster¶
Five nodes tolerate two failures and spread queue leaders over more hosts. Quorum queues still default to three members, so declare with x-quorum-initial-group-size: 5 (or use CMR with target_group_size = 5) if you want five replicas; more replicas mean more replication traffic per publish.
Kubernetes (Cluster Operator)¶
Use the RabbitMQ Cluster Operator to declare clusters with a RabbitmqCluster resource, and the Messaging Topology Operator to manage queues, exchanges, users and policies as YAML.
apiVersion: rabbitmq.com/v1beta1
kind: RabbitmqCluster
metadata:
name: prod
namespace: rabbit
spec:
replicas: 3
resources:
requests:
cpu: 4
memory: 8Gi
limits:
cpu: 4
memory: 8Gi
persistence:
storage: 200Gi
storageClassName: ssd
rabbitmq:
additionalConfig: |
default_queue_type = quorum
disk_free_limit.absolute = 4GB
anonymous_login_user = none
Note
Do not add cluster_partition_handling on 4.3+; it has no effect. The operator derives the memory watermark from the container limit, so set requests equal to limits for memory.
Blue-Green migration¶
For upgrades that cannot be done in place (3.12 to 4.x, 3.13 with experimental Khepri, or leaving classic mirrored queues behind), build a new cluster, import definitions, and move traffic with Shovels or Federation. Since 4.2, rabbitmqadmin v2 includes commands to automate parts of a 3.13 to 4.2 Blue-Green migration (Blue-Green guide).
Sizing¶
| Resource | Guidance |
|---|---|
| CPU | 4 cores minimum per node (official); quorum queues and TLS are CPU-heavy |
| Memory | 4 GiB minimum per node (official); keep at least 30% of RAM for the OS when raising the watermark above 0.7 |
| Disk | Fast local SSD/NVMe. Quorum queues fsync their WAL; streams are I/O-bound. Avoid NAS |
| Free disk alarm | disk_free_limit.absolute at least equal to the memory watermark (for example 4GB on an 8 GiB node) |
| Network | Replication multiplies publish traffic by the replica count; size inter-node links accordingly |
| File descriptors | Raise nofile well above the expected connection count (tens of thousands for busy nodes) |
Best Practices¶
- Default to quorum queues for durable work:
default_queue_type = quorumnode-wide, or--default-queue-type quorumper vhost. - Keep the delivery limit (default 20) and attach a dead-letter exchange via policy so poison messages are parked, not dropped.
- Use publisher confirms and consumer acknowledgements for anything that must not be lost.
- Set per-consumer prefetch (
basic.qos); global QoS is denied by default since 4.3 and streams reject it. - Use streams for large fan-out or replay instead of many fanout-bound queues.
- Federation or Shovel across WANs, not clustering; a cluster needs low-latency links.
- One vhost per environment or tenant with tight permission regexes.
- Enable the Prometheus plugin and scrape
:15692/metrics. - Long-lived connections, one channel per thread; opening a connection per request causes churn.
- Disable anonymous logins and keep
gueston loopback. - Keep large payloads out of messages: the default
max_message_sizeis 16 MiB since 4.0; store blobs elsewhere and send IDs.
Configure TLS¶
Enable TLS listeners for clients (and the same ssl_options apply to other listeners through their own prefixes, such as management.ssl.* and stream.listeners.ssl.default).
# rabbitmq.conf
listeners.tcp = none # close plaintext AMQP
listeners.ssl.default = 5671
ssl_options.cacertfile = /etc/rabbitmq/certs/ca.pem
ssl_options.certfile = /etc/rabbitmq/certs/server.pem
ssl_options.keyfile = /etc/rabbitmq/certs/server.key
ssl_options.verify = verify_peer
ssl_options.fail_if_no_peer_cert = true
ssl_options.versions.1 = tlsv1.3
ssl_options.versions.2 = tlsv1.2
Check it with rabbitmq-diagnostics listeners and openssl s_client -connect host:5671. Revocation checking (CRL/OCSP) needs separate configuration if you require it.
Encrypt inter-node traffic¶
Erlang distribution (25672) is plaintext by default. Point the runtime at the inet_tls distribution module (Inter-node TLS guide):
# rabbitmq-env.conf or environment
SERVER_ADDITIONAL_ERL_ARGS="-pa $ERL_SSL_PATH \
-proto_dist inet_tls \
-ssl_dist_opt server_certfile /etc/rabbitmq/certs/combined_keys.pem"
CLI tools need matching RABBITMQ_CTL_ERL_ARGS to reach the node.
Configure OAuth 2.0¶
The rabbitmq_auth_backend_oauth2 plugin validates JWTs (it does not do opaque-token introspection). Since 3.13 the simplest setup is to point it at the IdP issuer and let it discover the JWKS endpoint.
# rabbitmq.conf
auth_backends.1 = rabbit_auth_backend_oauth2
auth_backends.2 = rabbit_auth_backend_internal # optional fallback
auth_oauth2.resource_server_id = rabbitmq
auth_oauth2.issuer = https://idp.example.com/realms/prod
auth_oauth2.preferred_username_claims.1 = preferred_username
# map IdP-specific scopes to RabbitMQ scopes
auth_oauth2.scope_aliases.orders-producer = rabbitmq.write:prod/orders.* rabbitmq.configure:prod/orders.*
auth_oauth2.scope_aliases.orders-consumer = rabbitmq.read:prod/orders.*
Clients send the token as the password (username may be empty). AMQP 1.0 clients can refresh the token on a live connection (4.1+). The draft 4.4.0 release notes tighten OIDC discovery validation (HTTPS endpoints, matching issuer); plan for it if your IdP is non-compliant.
Set Up Users and Permissions¶
rabbitmqctl add_vhost prod --default-queue-type quorum
rabbitmqctl add_user orders-svc 'use-a-generated-secret'
rabbitmqctl set_permissions -p prod orders-svc '^orders\.' '^orders\.' '^orders\.'
rabbitmqctl set_topic_permissions -p prod orders-svc amq.topic '^orders\.' '^orders\.'
rabbitmqctl set_user_tags ops-alice monitoring
rabbitmqctl delete_user guest
The permission arguments are configure, write, read (regular expressions on resource names). See Reference: Access Control for what each permission and tag covers.
Enable Auditing and Observability¶
- Logs: by default in
/var/log/rabbitmq/; authentication failures and permission denials are logged atinfo/warning. - Events:
rabbitmq-plugins enable rabbitmq_event_exchangepublishes connection, channel, user, vhost and policy events toamq.rabbitmq.event; bind a queue and ship it to a SIEM. - Tracing:
rabbitmq_tracing(Firehose) records published and delivered messages; it is expensive, so enable it only while debugging. - Metrics:
rabbitmq-plugins enable rabbitmq_prometheus, then scrape:15692/metrics; the official Grafana dashboards live in therabbitmq-serverrepo underdeps/rabbitmq_prometheus/docker/grafana. After upgrading to 4.2+, update Raft dashboards (metric names changed). - Secrets in parameters: Federation and Shovel URIs contain credentials; restrict who can read runtime parameters (a 2026 advisory fixed disclosure to monitoring users).
Troubleshooting¶
Memory alarm: publishers blocked¶
Symptom: publishers receive connection.blocked; the management UI shows a memory alarm.
Causes: large backlogs, many connections/channels, big management stats retention, a watermark too high for the container.
Fixes:
rabbitmq-diagnostics memory_breakdownto find the largest consumer of memory.- Drain or shorten the backlog; move heavy queues to quorum queues or streams.
- Check the node sees the right memory limit (
total_memory_available_override_valuein containers). - Raise
vm_memory_high_watermark.relative(default 0.6) only as a temporary measure.
Disk free alarm¶
Reduce stream retention (max-age, max-length-bytes via policy), remove unused queues, or grow the volume. A quorum queue stuck in a requeue loop also grows its log; check delivery-limit.
Quorum queue has no leader¶
Symptom: rabbitmq-queues quorum_status NAME shows no leader; publishes time out.
Cause: fewer than a majority of members reachable (for example 1 of 3).
Fix: restore the missing nodes or network. If a node is gone permanently, remove it with rabbitmqctl forget_cluster_node (4.3 removes its queue and stream members first) or rabbitmq-queues delete_member, then add a member on a healthy node. force_reset is deprecated since 4.1 and incompatible with Khepri.
Metadata operations fail during a partition¶
Since 4.3 (and on 4.2 clusters with Khepri), declaring queues, bindings or users fails on the minority side of a partition. This is expected. Check with rabbitmqctl cluster_status and restore connectivity; there is no partition-handling strategy to tune.
Slow consumers¶
Low consumer utilisation with a high unacked count means consumers are the bottleneck: add consumers, lower per-message work, or tune prefetch. On 4.3 quorum queues, consumer-timeout returns messages held too long.
Node.js clients cannot connect after upgrading to 4.1+¶
amqplib older than 0.10.7 uses a 4096-byte pre-auth frame_max; 4.1 requires at least 8192. Upgrade the library.
MQTT or Web MQTT clients disconnecting¶
- Confirm
rabbitmq_mqtt/rabbitmq_web_mqttare enabled on every node and listeners are configured (mqtt.listeners.tcp.default,mqtt.listeners.ssl.default). - Since 4.1 the default MQTT max packet size is 16 MiB; larger packets are rejected.
- Since 4.0
mqtt.default_useris gone; useanonymous_login_useror real credentials.
Upgrade to 4.3¶
flowchart TD
Now{"Current series?"} -->|"3.12.x or older"| BG["Blue-Green migration<br/>to a new 4.2+ cluster"]
Now -->|"3.13.x with experimental Khepri"| BG
Now -->|"3.13.x (Mnesia)"| To42["Enable all feature flags,<br/>rolling upgrade to latest 4.2.x"]
Now -->|"4.0.x or 4.1.x"| To42
Now -->|"4.2.x"| Khepri["rabbitmqctl enable_feature_flag all<br/>(includes khepri_db)"]
To42 --> Khepri
Khepri --> Erl{"Erlang 27+ on every node?"}
Erl -->|no| UpErl["Upgrade Erlang first"]
UpErl --> Roll
Erl -->|yes| Roll["Rolling upgrade node by node<br/>to latest 4.3.x"]
Roll --> Clean["Remove cluster_partition_handling keys,<br/>check deprecated features,<br/>rabbitmq-queues rebalance all"]
- Read the 4.3.0 release notes. Only 4.2.x upgrades in place to 4.3.x.
- Check applications for features now denied by default: transient non-exclusive queues, global QoS,
x-queue-mode, AMQP 1.0 v1 addresses. Opt back in temporarily withdeprecated_features.permit.<name> = trueon all nodes before upgrading. - Enable every stable feature flag:
rabbitmqctl enable_feature_flag all, then confirmkhepri_dbis enabled withrabbitmqctl list_feature_flags. - For each node:
rabbitmq-diagnostics check_if_node_is_quorum_critical(must pass),rabbitmq-upgrade drain, stop, upgrade the package, start,rabbitmq-upgrade revive, thenrabbitmq-diagnostics check_runningbefore moving on. Automation can userabbitmq-upgrade await_online_quorum_plus_one. - Keep the mixed-version window to hours, not days.
- Afterwards run
rabbitmq-queues rebalance alland remove partition-handling keys fromrabbitmq.conf.
Migrating off classic mirrored queues
After an upgrade to 4.x, mirroring policies silently stop replicating. Declare replacement quorum queues (queue type is immutable), move consumers, and use a Shovel to drain the old classic queues. See Migrate mirrored classic queues to quorum queues.
Cost Analysis¶
| Cost | Driver |
|---|---|
| Compute | Erlang VM baseline is moderate; quorum queues and TLS add CPU per message |
| Storage | Quorum queue WAL fsyncs and stream segments; provision SSD/NVMe and headroom above disk_free_limit |
| Memory | Connections, channels and quorum queue in-memory index; 4.3 roughly halves per-message overhead in many cases |
| Network egress | Replication traffic within the cluster; Federation/Shovel duplicate traffic across regions |
| Tanzu RabbitMQ | Commercial license from Broadcom; required for patches to out-of-community-support series (pricing not public) |
| Managed services | CloudAMQP (per plan), Amazon MQ for RabbitMQ (per broker-hour and storage) |
Commands & Recipes¶
Bootstrap and cluster¶
# On node 1
rabbitmq-diagnostics status
# On node 2: join node 1's cluster
rabbitmqctl stop_app
rabbitmqctl reset
rabbitmqctl join_cluster rabbit@node1
rabbitmqctl start_app
rabbitmqctl cluster_status
Declare a quorum queue, exchange and binding (rabbitmqadmin v2)¶
rabbitmqadmin --vhost prod queues declare --name orders --type quorum --durable true
rabbitmqadmin --vhost prod exchanges declare --name orders.x --type topic --durable true
rabbitmqadmin --vhost prod bindings declare --source orders.x --destination-type queue \
--destination orders --routing-key 'orders.created.*'
# delivery limit, length cap and DLX via policy (preferred over queue arguments)
rabbitmqadmin --vhost prod policies declare --name orders-limits --pattern '^orders$' \
--definition '{"delivery-limit":5,"max-length":1000000,"dead-letter-exchange":"orders.dlx"}' \
--priority 1 --apply-to quorum_queues
Declare a stream and a super stream¶
rabbitmqadmin --vhost prod streams declare --name events --expiration 7D
rabbitmqctl set_policy -p prod events-retention '^events$' \
'{"max-length-bytes":50000000000}' --apply-to streams
rabbitmq-streams add_super_stream invoices --partitions 3 --vhost prod
Federation upstream¶
rabbitmqctl set_parameter -p prod federation-upstream us-prod \
'{"uri":"amqps://fed-user:secret@us-prod.example.com:5671"}'
rabbitmqctl set_policy -p prod federate-orders '^orders\.' \
'{"federation-upstream-set":"all"}' --apply-to exchanges
Diagnostics¶
rabbitmq-diagnostics status
rabbitmq-diagnostics memory_breakdown
rabbitmq-diagnostics check_alarms
rabbitmq-diagnostics check_running
rabbitmq-diagnostics observer # text-mode observer (top-like)
rabbitmqctl list_queues name type messages messages_ready consumers
rabbitmqctl list_connections user host channels state
rabbitmq-queues quorum_status orders
rabbitmq-streams stream_status events
rabbitmqadmin health_check local_alarms
Prometheus¶
perf-test¶
docker run -it --rm pivotalrabbitmq/perf-test:latest \
--uri amqp://app:pwd@host:5672 \
--producers 10 --consumers 10 \
--rate 10000 --confirm 100 \
--queue-pattern 'q-%d' --queue-pattern-from 1 --queue-pattern-to 50 \
--quorum-queue
Cross-references¶
- Explanation: queue types, Khepri and the security model behind these recipes.
- Reference: config defaults, ports, upgrade paths, hardening checklist.
- Messaging: comparisons with Kafka, NATS, Redpanda and Pulsar.