A spend cap you haven't driven to zero is a number in a config file
We price in prepaid credits so the worst case is meant to be boring. You hit zero, generation stops, nobody gets a bill they didn't agree to.
Then we shipped a spend function that was writing bad balance events. The ledger drifted from what people had actually spent, so the ceiling I thought I had wasn't the ceiling. Nothing failed, nothing alerted, and the number in the config was still correct the whole time.
The part I think is getting worse rather than better: the diff was fine. Reading it, there was nothing to catch. The failure lived in a state nobody wrote a test for, and review is good at bad code and bad at missing cases. Shipping got fast, the sitting and thinking about what happens at the boundary didn't get any faster.
So the rule now is that anything touching balance gets driven to zero and past it before it merges. Not read, driven.
If you've got a hard cap in your product, has anything in your test suite actually hit it?
Replies
Same with agent guardrails. Reverting the fix still left the test green because the assertion was never testing the bug. If you haven't watched the check fail with your own eyes, it is just decorative code.
@konstantin_tikhaev Decorative is exactly the right word. Mine was checking that the cap was configured, not that it held at the boundary, and in a diff those two read identical. Watching it go red first is the step I skipped, because green felt like it was telling me the same thing.
@asadmalik901 Diffs make schema checks look like behavior checks. If you don't break the runtime logic and verify the failure, you are just testing the parser.
The test that worries me is the one that passes because it checks the policy exists, not because it checks the policy was enforced at the number that matters. I hit something close to this with a Stripe metered billing check. The test asserted the webhook handler ran, not that it ran before the next request was allowed through. A slow webhook and a fast second request could both pass the suite and still let someone through a window where the cap had technically fired but had not yet taken effect. The fix that mattered was a test that fires two requests in quick succession and checks the second one died, not one that checks the first one logged a cap event. Reading a cap check feels like reading a boundary check, and it rarely is.