Unit tests passed. Coverage looked fine. The first integration test failed because the handler needed a database, the database needed a migration, the migration conflicted with the in-memory provider someone assumed would work, and the fake message bus did not implement the interface the outbox dispatcher actually uses.
That failure was not a testing problem. It was an architecture problem we had not looked at yet.
Unit tests optimize for isolation; systems optimize for wiring
A unit test proves that a piece of logic behaves correctly when its dependencies are replaced with mocks. That is useful. It is also a controlled fiction. The real dependencies have lifetimes, transaction boundaries, ordering constraints, and failure modes that mocks flatten away.
Integration tests force the question: can these pieces actually run together? Does the repository save what the handler thinks it saved? Does the consumer see the message after the transaction commits? Does the API return the error shape the client expects when the database is slow?
I have seen teams with thousands of unit tests and twelve integration tests treat the latter as “slow and flaky” while the production system failed on wiring nobody had exercised.
What integration tests expose that units hide
Transaction boundaries. A handler that publishes inside a request and commits after will pass unit tests with a mocked bus. An integration test against a real database shows the consumer reading a message for a row that does not exist yet.
Schema and mapping drift. EF configurations, DTO mappings, and API contracts diverge quietly. A test that hits the HTTP endpoint and reads back from the database catches mismatches that isolated mapper tests miss.
Configuration. Connection strings, feature flags, options binding — all of that is invisible to unit tests and very visible at 2 a.m. when staging works and production does not because one environment variable differs.
Realistic failure paths. Timeouts, deadlocks, unique constraint violations. Mocks usually throw exactly what you told them to throw. Production throws what it throws.
The cost is real; skipping it is more expensive
Integration tests are slower. They need containers or a shared test database. They fail for environmental reasons. Flaky tests erode trust fast.
The answer is not to skip them. It is to be deliberate about what they prove. I do not need an integration test for every branch of a pricing function. I do need one for each critical path that crosses a boundary I care about in production: write path, message publish, external API call with retry, auth.
A small set of integration tests that cover the spine of the system beats a large set that only touches the edges in isolation.
How I think about the split now
Unit tests for rules, calculations, state transitions inside a single type. Fast feedback, easy to run on every commit.
Integration tests for use cases: place order, process payment webhook, expire a subscription. One test per happy path, plus the failure modes that have burned us before.
End-to-end tests sparingly, for flows the business names in incidents.
When an integration test is hard to write, I treat that as signal. Maybe the handler does too much. Maybe the database is doing work that should be explicit. Maybe the system has so many moving parts that nobody can describe the path from request to side effect without a whiteboard.
That discomfort is the test doing its job.