I used to treat the outbox as something you add later. First you get the feature out. Then you “make messaging reliable”. That order is wrong.
The first time I saw it break, the business write succeeded and the publish did not. The user had an order. The warehouse never heard about it. The next time it broke the other way: the message went out, the transaction rolled back, and a consumer billed a customer for a sale that no longer existed.
Those are not edge cases. They are the default shape of a dual write.
Two commits are not one decision
A database commit and a broker publish are two different durability stories. Between them you get process death, a timeout that lied, a broker that accepted the message and then dropped the connection, a retry that ran after the request already looked successful.
If you write to SQL and then publish, you can lose the event. If you publish and then write, you can emit an event for a state that never landed. Retrying the HTTP request does not reconcile those two worlds. It usually makes them worse.
This is the part teams skip in design reviews. They draw a box labeled “publish event” after “save order” and treat the arrow as atomic. It is not. The failure is not theoretical. It shows up the first time a deploy lands in the middle of a checkout, or a pod is evicted, or the broker is slow enough for the request to time out after the message was already accepted.
Consumers cannot fix a dishonest producer
Idempotent consumers are necessary. They are not sufficient.
An inbox, a unique message id, a “process this payment at most once” guard — all of that assumes the producer is telling a consistent story. If the producer can emit “OrderPlaced” without a matching row, the consumer will faithfully create a shipment for nothing. If the producer can commit the row and never emit, the consumer has nothing to be idempotent about.
Reliability at the edge does not repair a hole in the middle.
What the outbox actually buys
The outbox is a boring idea: the business change and the intent to publish are the same commit. The row in Orders and the row in OutboxMessages either both exist or neither does. Something else — a poller, CDC, a background dispatcher — later turns that intent into a broker message.
You still get duplicates. The dispatcher can publish and crash before marking the outbox row as sent. The broker can deliver twice. That is fine. Duplicates are a known, local problem. Missing events and phantom events are a global one.
I care about that distinction more than I care about elegance. A consumer that sees the same OrderPlaced twice can no-op. A system that has an order without an event has a silent operational hole. You find it days later, when someone asks why the warehouse is empty.
The cost you are actually paying
Outbox is not free. You take on a publisher process, at-least-once delivery, and the need to think about ordering if your domain cares about it. You may need CDC if polling a table at your write volume is ugly. You still need idempotency on the other side.
That cost is real. It is also smaller than the cost of a reconciliation job you invent after the first incident, plus the Slack thread where nobody can prove whether the event was sent.
The version I regret is the one that starts with “we’ll publish in the handler and add an outbox if we need it”. By the time you need it, you already have two sources of truth and a week of forensic SQL.
When I would still skip it
If there is no broker, skip it. If the write and the side effect live in the same transaction — an email queued in the same database, a projection in the same SQL commit — you do not have a dual-write problem.
If the event is a nice-to-have signal and the system can be wrong for a while, say so. Do not dress that up as “event-driven architecture”.
What I no longer do is ship a production write path that mutates state and publishes to a bus as two independent I/O calls, then call the gap a future improvement. That gap is the design.