Messaging is sold as a way to decouple services. What it actually decouples is time. That is a real benefit. It is also the source of most of the pain.
A synchronous call fails in the request. You see it. The user sees it. You retry or you show an error. An asynchronous message fails in a worker you were not looking at, in a topology nobody fully drew, with a payload that may already have been retried twice.
I have introduced queues to make a system “more scalable” and made it harder to change, harder to debug, and no more scalable than the database that was already the bottleneck.
What you buy, and what you take on
You buy the right to finish the user’s request before the rest of the work happens. You buy a boundary that can be down without taking the writer down with it. You buy a place to absorb a burst.
You take on at-least-once delivery, out-of-order arrival, poison messages, and the need to version a contract that now lives outside the process. You take on replay. You take on the question “what is the current state?” when the answer is spread across a log and three consumers that have not caught up.
None of that is a reason to avoid messaging. It is a reason not to reach for it because a diagram looked cleaner with arrows between boxes.
The failure modes were not in the old design
A checkout that writes an order and then calls billing over HTTP has a simple failure story. The call worked, or it did not. Timeouts are ugly, but they are local.
Put a queue in the middle and you inherit a new set of states:
- The message was published and never consumed.
- The message was consumed and the handler crashed after the side effect.
- The message arrived twice, a week apart, because someone replayed a subscription.
- The message arrived before the one it depends on.
- The payload is valid JSON for last quarter’s schema and invalid for this week’s consumer.
The old system did not need a poison-message strategy. The new one does, whether you wrote it down or not. If you did not, the strategy is “the ops person restarts the worker and hopes”.
Operational cost shows up after the happy path
The first demo is always fine. Publish, consume, green test. The cost arrives when you need to know why customer 4812 did not get a confirmation, and the only trail is a message id in one system and a row in another, with no shared timeline.
It arrives when a consumer must be redeployed because a field was renamed, and you discover that three older producers are still emitting the previous shape. It arrives when someone asks for a replay of Tuesday and you realize replay is a product feature you never built.
I now treat “we will add observability later” as a decision to fly without instruments. Messaging without correlation ids, without a way to see lag, without a defined replay path, is not architecture. It is a delayed incident.
A queue is not an architecture
I have seen teams split a modular monolith into messages between in-process components that still share a database and a release. Nothing got more independent. Latency got worse. The debugger got less useful.
Messaging earns its place when a part of the work can and should happen later, or when two teams need a contract that can survive independent deploys. If both sides still ship together, and the reader needs the write to have happened just now, a queue is an expensive function call.
Sometimes the honest design is: write the row, call the next thing, live with the coupling. You can extract a message later if the temporal boundary becomes real. Extracting a queue first, hoping independence will follow, usually gives you a distributed monolith with worse failure modes.
When I reach for it anyway
I want messaging when the user should not wait for the side effect, and the side effect can be wrong for a short time without becoming a support ticket. I want it when a downstream system has a different availability story than the writer. I want it when the volume of work is bursty and the writer should not absorb that burst on its request thread.
I do not want it as a default communication style between any two boxes on a slide.
The question I ask now is not “should this be event-driven?”. It is “which failure modes are we buying, and did we need them?”. If the answer is that we needed decoupling in time, the queue stays. If the answer is that we wanted the architecture to look grown-up, a synchronous path would have been the smaller system.