Self-Healing Beats Retrying

Three pipelines, three kinds of failure, and none of them were fixed by a better retry loop. Recoverable by resuming, recoverable by giving up early, and not recoverable at all need three different patterns.

I built three pipelines this year that each had to survive partial failure, and every time, my first instinct was “just add a retry.” Every time, that instinct was wrong, because a blind retry treats every failure as the same failure, and they’re not.

Take a batch job generating copy for a few thousand catalog items, two model calls per item, real money per call. Crash at item 3,000 of 8,000, and a naive retry-from-scratch either burns 3,000 items of spend twice or, worse, silently skips ahead and nobody notices the first 3,000 were never checked. Neither is acceptable. A JSON-lines checkpoint fixes it: one line per completed item, written append-and-fsync as it goes, so a restart skips everything already marked done and picks up exactly where it stopped. A crash costs you the one in-flight item, not the run. The detail that took a second pass to get right: a crash mid-write leaves a torn final line in that file, and if you don’t handle it explicitly, the file fails to parse on reload and you lose the whole log instead of just the last item. Detect the torn line, discard it, keep everything before it. Skip that step and “resumable” quietly becomes “corruptible.”

A different shape of problem shows up in an approval queue, where an LLM proposes actions and a human approves or rejects them before a dispatcher executes the approved ones. Here retrying doesn’t lose you progress, it wastes it: a permanently broken handler keeps re-entering the execution queue on every cycle, and the noise buries the real backlog underneath it. What this needed wasn’t resumability, it was a ceiling: three attempts, and on the third failure the proposal auto-rejects with the actual error attached to the decision, instead of retrying a fourth time against something that was never going to work.

A report-generation API taught me failures split further still. Request, poll, download. Some statuses are transient (still generating, poll again), some are terminal (failed, and no amount of polling changes that). Treat them the same and you poll a dead job for twenty minutes before giving up, burning your deadline on a status that told you the answer the moment it arrived. Stopping immediately on a terminal status means recognizing, on sight, that trying again cannot help, and that only works if the code distinguishes the two statuses in the first place instead of lumping every non-200 into “try later.”

Three pipelines, three different fixes, and none of them came from writing a better retry loop. They came from stopping to ask what kind of failure I was actually looking at: recoverable by resuming, recoverable by giving up early, or not recoverable at all.

The checkpoint-and-resume pattern lives in catalog-copy. The failure-cap pattern is in approval-loop’s dispatcher. The transient-versus-terminal distinction is in settlement-pl’s report poller. Next time you’re about to wrap something in a retry, figure out which of the three you’re dealing with first.