Overview
How might we let agents recover from failure without silent retries, mystery fallbacks, or endless loops?
When to use
- LangChain, Zapier, and Make workflows where tool calls, APIs, or models fail intermittently.
- Long-running agents like AutoGPT that need backoff, alternate paths, or re-planning within a budget.
- Products where users configure retry limits, fallback behavior, and escalation instead of one-size-fits-all errors.
- Any surface that pairs with failure disclosure: name the miss first, then show how recovery is trying to fix it.
When to skip
- Simple, single-step tools where a failure just needs a plain error message and a manual retry button.
- Actions with real-world side effects where automatic retry could duplicate an already-completed effect (like a payment).
- Early-stage products where exposing granular retry configuration adds settings surface nobody will tune.
States
Design the recovery policy and its visible timeline, not only a generic error toast.
- 01
Failed
A tool, model, or connector errors. The miss is named and recovery policy activates.
- 02
Retrying
The agent retries within a configured cap. Attempt count and backoff timing stay visible.
- 03
Fallback
Retries exhaust. A defined alternate runs: degraded mode, queued retry, or substitute workflow.
- 04
Escalated
Automatic recovery stops. The person or an operator gets handoff, approval, or manual intervention.
- 05
Recovered
A retry or fallback succeeds. Status returns to normal and the timeline records what worked.
- 06
Configured
The user sets retry maximum, fallback strategy, and escalation threshold before or between runs.
Key UX elements
The parts that must be present for configurable recovery to feel reliable.
Retry
Cap attempts and show the count.
Display attempt N of M so people know the agent is still trying, not stuck or silently looping.
Backoff
Make delay between tries legible.
Name the wait or queue window. Immediate hammering on a dead endpoint erodes trust faster than a pause.
Fallback
Name the alternate path plainly.
Degraded mode, queued retry, or substitute workflow must read as an intentional switch, not a hidden downgrade.
Escalate
Stop with a final handoff.
When the budget is gone, offer human review, approval, or a clear dead-end, never infinite retry.
Timeline
Log recovery events in order.
A running event list beats a single spinner. People should see detect, retry, fallback, and resolve.
Policy
Let operators tune before the next incident.
Retry max, fallback choice, and escalation sensitivity belong in settings people can adjust without redeploying code.
Rules
Silent infinite retries that burn time or budget with no visible cap or escalation.
Fallback actions that change behavior without telling the user a fallback was used.
Treating every failure the same way regardless of whether it is transient or permanent.
No final escalation path, leaving the agent stuck retrying forever after repeated failures.
Evidence
| Product | Implementation |
|---|---|
| LangChain | Retry and fallback chains reroute to alternate tools or models after failures. |
| AutoGPT | Failed actions trigger re-planning or alternate approaches within budget limits. |
| Zapier | Configurable retry rules and error paths reroute failed automation steps. |
| Make (Integromat) | Error handlers define retry, ignore, or fallback routes per workflow module. |