Six checks passed, but did we actually test failure recovery?
by•
We ran a Lovable checkout app through FetchSandbox. Six checks passed. Two stayed not checked.
We delayed an email response, but the webhook still returned success. We hadn't shown it recovering after a failed callback. So the overall run stayed incomplete.
That's the thing I want us to get right as we build this. Did the test cause the failure, or did the app just succeed? Same question for a background job that might crash halfway through.
What evidence do you ask for before trusting a recovery test?
82 views


Replies
I think a passing test that never caused the failure is worse than no test. It gives you false comfort. I got burned once when a retry "worked" only because the first call never actually failed.
@david_turner12 Yeah David, same thing showed up in our run. The email response was delayed but the webhook still came back successful, so we never actually put failed callback recovery under any pressure.
we hit almost the exact version of this with a callback timing bug in our voice agent - a test suite "passed" the failed-callback path for months because the mock retried instantly, so the retry logic never had to survive an actual delay. the bug only showed up once a real customer complained their callback landed late and got dropped. what I ask for now is a timestamp diff: show me the gap between when the failure was supposed to happen and when the recovery code actually started running. if that gap is near-zero or suspiciously clean, I don't trust the pass, a real failure is messy and has jitter
@galdayan The timestamp gap idea is solid. We need to make the actual failure time and retry start clearer on the receipt instead of just showing the planned delay window. Did your late callback arrive after the job was already marked complete, or did it get discarded mid-flight?
@rnagulapalle after - the job's state machine had already moved on, so the late callback just got silently dropped at the door, no error, no retry, nothing in the logs that screamed "this was late." that's actually what made it worse than a hard failure, a crash would've paged someone. the fact it just quietly no-opped is why it sat undetected for months. that's the receipt problem too I think - "discarded because too late" and "recovered successfully" need to look completely different on the output, not just a shared green checkmark
@galdayan Yeah, a callback that arrives after completion should not register as recovery. The receipt needs the job state at arrival and whether that callback was actually used or quietly dropped.
@rnagulapalle agreed, and that's a cleaner split than what we actually shipped - we just added a "used" boolean next to the job state, so the receipt now answers both questions at once instead of making someone infer it from timestamps. small change, but it turned a 10 minute debugging session into a 10 second glance at the receipt
I would add a crash after the side effect but before the job records success
Then restart the worker and check that the order or email is not created twice
Recovering the job is only half the check
@al_kub This delayed-response run does not prove crash-and-restart recovery. The real check is a separate test: crash after the order or email is created, restart the worker, and verify the original still exists and no second one was created.