If all your tests pass, how do you know the running system is actually right?
Spent yesterday adding a new API. 54 tests green. No dangling refs. Then I actually called it.
Four things wrong, none of them raised, none logged, every test stayed green.
What broke
Every read-back 404'd because I named the endpoint after the resource instead of copying the path from the spec. An ID pattern used a wildcard the generator doesn't recognize, so it handed callers the literal string ent_%%%%%%%%. Seeded records and runtime records had different ID formats because seeding ran outside the context that tracks them. One endpoint returned a list of lists because a sub-resource got seeded from the envelope instead of the record inside it.
The one that bothered me most
Matched my exact symptom strings at 0.95 confidence. Paraphrased the same bug like a person would type it and got 0.4. It looked deep. It was keyed.
The actual question
How do you catch config bugs like this? Not code bugs. Every layer right on its own, running system still wrong. Green tells you what you meant to build, not what it does.
Launching what this led to tomorrow if you want to see where it ended up.

Replies
This is the class of bug that reaches me before it reaches the devs, nothing errors, and then a customer answers our welcome email asking why the link goes to a plan that doesn't exist. Nobody's test owned that seam, the email tool and the billing config were each right on their own. The only thing that's ever caught these for me is boring routine: once a week I walk the real funnel like a stranger, real signup, real email, real purchase, because the seams are exactly where green can't see.