Almost every distributed system reaches a point where two components need to talk without waiting for each other, and the answer is some form of asynchronous messaging. What follows is usually a debate about products, RabbitMQ or Kafka, Service Bus or Event Hubs, as if the question were which is best. It is not. The products fall into two paradigms that look similar and behave completely differently, and the single most useful thing you can do before choosing a tool is work out which paradigm your problem actually is. Get that right and the shortlist writes itself. Get it wrong and you will spend the life of the system fighting a platform that was never built for what you are asking of it.
This is a practitioner's tour of the landscape: the distinction that matters, the platforms in each camp including the ones people forget, the mechanics that will bite you regardless of which you pick, and the native choices in AWS and Azure.
The one distinction that decides everything
Here is the whole article in two sentences. A message queue is a post office: a message is an instruction, it is routed to a worker, the worker acknowledges it, and it is deleted. An event log is a journal: an event is a fact that happened, it is appended to a durable log and retained, and many independent consumers can read it, at their own pace, and rewind to re-read the past. That difference, delivered-then-deleted versus appended-and-retained, drives everything else: how you scale, how you route, how many consumers you can have, and whether you can ever replay history.
The tell is in the payload. If the thing you are sending is an instruction, do this piece of work once, send this email, resize this image, take this payment, you want a queue, because once the work is done the instruction has no further value. If the thing you are sending is a fact, an order was placed, a user logged in, a sensor read 21 degrees, you may want a stream, because that fact can matter to several parts of the system, now and later, and losing it after the first reader would be a mistake. Instructions empty out of a queue. Facts accumulate in a log.
The message queues, and where each earns its place
RabbitMQ is the reference message broker. Its strength is flexible routing through exchanges and bindings, so you can send a message to exactly the right queues based on content or headers, and per-message acknowledgement with dead-lettering built in. It speaks several protocols, AMQP, MQTT and STOMP, which makes it a natural fit for mixed clients including IoT. It is comparatively easy to run, a single Erlang binary with a genuinely useful management UI. Modern versions added quorum queues for stronger durability and even a native streams feature that borrows the retention idea, but its heart is still the classic task queue. Reach for it when you need sophisticated routing, RPC-style request and reply, priorities, and reliable per-message delivery at moderate volume.
Azure Service Bus is the enterprise broker of the Azure world, and the natural queue choice if you are already there. It offers queues and topics with subscriptions for publish and subscribe, sessions for ordered processing, dead-letter queues, scheduled and deferred messages, duplicate detection and transactions. It is fully managed, so you are not operating a broker, and it integrates cleanly with the rest of Azure. Reach for it for line-of-business workflows, coordination between services, and anywhere you want enterprise messaging features without running the infrastructure yourself.
Amazon SQS is the queue you reach for on AWS when you want the least possible operational burden. It is fully managed, effectively infinitely scalable, and cheap, with standard queues for maximum throughput and FIFO queues when you need ordering and exactly-once processing within the queue. It is deliberately simple: no complex routing, which is what SNS (fan-out publish and subscribe) and EventBridge (event routing) are for alongside it. Reach for SQS for straightforward asynchronous work distribution on AWS where simplicity and zero operations matter more than routing sophistication.
Worth knowing beyond the big three: ActiveMQ, and its managed form Amazon MQ, when you need a standards-based broker (JMS, AMQP) often to migrate an existing Java or enterprise system without rewriting it; IBM MQ, still the backbone of a great deal of banking and enterprise integration, where it is chosen for pedigree and guarantees rather than novelty; and NATS with its JetStream persistence, a lightweight, very fast, cloud-native option popular in Kubernetes and microservice environments.
The event streams, and what the log buys you
Apache Kafka is the reference event streaming platform and, for high-volume event-driven systems, close to a default. It is a distributed, partitioned, append-only log: events are retained for a configured time or size, consumers track their own position, and a single cluster comfortably handles enormous throughput. That retention is the superpower, it is what lets you replay history to backfill a new service, feed several independent consumers the same events, and use the stream itself as a system of record in event sourcing. It also carries a real operational cost: partitions, consumer groups and cluster management take expertise, though modern Kafka has removed the old ZooKeeper dependency in favour of its own KRaft consensus. Reach for Kafka when throughput is high, when events must be retained and replayable, when many teams consume the same data, or when you are building stream processing with Kafka Streams or Flink.
Azure Event Hubs is the Azure answer to streaming, and importantly it exposes a Kafka-compatible endpoint, so Kafka clients can often point at it with minimal change while Azure runs the platform for you. It is built for massive ingestion, telemetry, logs and event pipelines, and integrates with the Azure analytics stack. Reach for it when you want Kafka-style streaming on Azure without operating a Kafka cluster. Its sibling Event Grid is a different thing worth not confusing with it: a reactive event-routing service for discrete events (a blob was created, a resource changed), not a high-throughput stream.
Amazon Kinesis is the AWS-native streaming service for real-time data at scale, and Amazon MSK is managed Kafka for teams that specifically want Kafka on AWS without running it. Beyond the hyperscalers, Apache Pulsar separates serving from storage and has strong multi-tenant and geo-replication stories, which appeal at large scale; Redpanda is Kafka-compatible but written without the JVM for lower latency and simpler operations; and Redis Streams is a pragmatic middle option when you already run Redis and need lightweight streaming without standing up Kafka. Google Cloud Pub/Sub sits across the line as a scalable, fully managed hybrid that behaves like both for many use cases.
The native picks in your cloud
If you are on AWS or Azure, the honest default is to start with the managed native service and only reach for self-hosted Kafka or RabbitMQ when you have a reason the native option cannot meet, because the operational saving is large and real.
| You need | On AWS | On Azure |
|---|---|---|
| Simple async task queue | SQS | Service Bus (or Storage queues for basic needs) |
| Publish and subscribe fan-out | SNS | Service Bus topics |
| Event routing between services | EventBridge | Event Grid |
| Enterprise messaging features | Amazon MQ (or SQS FIFO) | Service Bus |
| High-throughput event streaming | Kinesis, or MSK for Kafka | Event Hubs (Kafka-compatible) |
| Full Kafka, managed | Amazon MSK | Event Hubs, or Kafka on AKS |
The mechanics that bite regardless of the tool
Whichever paradigm and product you choose, a handful of realities catch teams out, and they are worth designing for from the start rather than discovering in an incident.
Delivery is at-least-once, not exactly-once. Most systems guarantee that a message arrives at least once, which means it can arrive twice, on a retry after a consumer crashed mid-processing, for example. Some platforms advertise exactly-once, but it is narrow, conditional and easy to break the moment your processing touches an external system. The safe design is to assume at-least-once and make your consumers idempotent, so that processing the same message twice does no harm. Idempotency, not a vendor promise, is what actually protects you.
Ordering is local, not global. Order is guaranteed within a single queue or a single partition, and not across the whole topic. If your business logic depends on strict global ordering, you have to design for it, usually by routing related messages to the same partition through a partition key, and you have to accept the throughput limits that imposes. Assuming global order and not getting it produces bugs that are maddening to reproduce.
Poison messages need a dead-letter path. Sooner or later a message arrives that a consumer cannot process, malformed, referencing something that no longer exists, or simply triggering a bug. Without a dead-letter queue to move it aside after a few failed attempts, it sits at the head of the queue and blocks everything behind it, or loops forever consuming resources. A dead-letter queue and a retry policy are not advanced features to add later; they are part of a correct design.
Backpressure and operational burden are real costs. A stream absorbing a traffic spike is a feature; a queue growing without bound because consumers cannot keep up is an outage in waiting. And the platform with the most power, Kafka, carries the most operational weight. Choosing self-hosted Kafka without the team to run partitions, rebalancing and cluster health is one of the most common ways a good architecture becomes a bad on-call rota. If you cannot staff it, choose the managed equivalent.
The honest summary
Messaging is not a contest between RabbitMQ and Kafka, and it is not really about products at all until the last step. It is about recognising whether your payload is an instruction to be done once or a fact that many parts of the system may care about, now and later. The first is a queue: RabbitMQ, Azure Service Bus, Amazon SQS, and their kin. The second is a log: Kafka, Azure Event Hubs, Kinesis, and theirs. Name the paradigm, then pick the product that fits your cloud and your team's ability to operate it, defaulting to managed unless you have a clear reason not to.
And whichever you choose, design from day one for the three things that bite everyone: assume at-least-once delivery and make consumers idempotent, do not assume ordering you were not promised, and always have a dead-letter path. Most messaging problems in production are not the product's fault; they are one of those three assumptions, made silently, and found the hard way. Mature systems, incidentally, usually run more than one of these, a stream for events, a queue for tasks, because the two paradigms are complements, not competitors.
If you are designing an event-driven architecture and want the paradigm, the platform choice and the delivery guarantees decided so they hold up under load, that is the kind of work I do through Cyber Spartans.