Six checks passed, but did we actually test failure recovery?
by•
We ran a Lovable checkout app through FetchSandbox. Six checks passed. Two stayed not checked.
We delayed an email response, but the webhook still returned success. We hadn't shown it recovering after a failed callback. So the overall run stayed incomplete.
That's the thing I want us to get right as we build this. Did the test cause the failure, or did the app just succeed? Same question for a background job that might crash halfway through.
What evidence do you ask for before trusting a recovery test?
52 views


Replies
I think a passing test that never caused the failure is worse than no test. It gives you false comfort. I got burned once when a retry "worked" only because the first call never actually failed.
@david_turner12Â Yeah David, same thing showed up in our run. The email response was delayed but the webhook still came back successful, so we never actually put failed callback recovery under any pressure.
The recovery test that worries me is the one that passes because the retry fired fast enough to beat the failure window, not because it actually recovered. I built deadline tracking into ComplaintForge, so a letter's escalation trigger depends on a clock surviving a delay rather than succeeding despite one. What convinced me the tests were real was forcing the delay past the actual statutory window rather than past some arbitrary short timeout, because a fast retry hides exactly the gap you are trying to prove the system survives. The evidence I would ask for here is close to what Gal described: not that the check passed, but the wall clock time between the failure and the system noticing it, with proof that gap was long enough to be the real failure rather than a convenient one.
@oshylabs Do you test that full deadline with a controlled clock or just wait through it?
Good that it didn't show a green result just because nothing broke. If the failure never really happened, a pass doesn't tell you much.
@marina_gomel5Â Exactly, Marina. We never saw those failure scenarios actually trigger in the run, so we couldn't call those two tests passed.
we hit almost the exact version of this with a callback timing bug in our voice agent - a test suite "passed" the failed-callback path for months because the mock retried instantly, so the retry logic never had to survive an actual delay. the bug only showed up once a real customer complained their callback landed late and got dropped. what I ask for now is a timestamp diff: show me the gap between when the failure was supposed to happen and when the recovery code actually started running. if that gap is near-zero or suspiciously clean, I don't trust the pass, a real failure is messy and has jitter
@galdayan The timestamp gap idea is solid. We need to make the actual failure time and retry start clearer on the receipt instead of just showing the planned delay window. Did your late callback arrive after the job was already marked complete, or did it get discarded mid-flight?
@rnagulapalle after - the job's state machine had already moved on, so the late callback just got silently dropped at the door, no error, no retry, nothing in the logs that screamed "this was late." that's actually what made it worse than a hard failure, a crash would've paged someone. the fact it just quietly no-opped is why it sat undetected for months. that's the receipt problem too I think - "discarded because too late" and "recovered successfully" need to look completely different on the output, not just a shared green checkmark