Execute tasks reliably under transient failures without duplicating side effects or hiding risk. Use when network calls, package operations, APIs, locks, rate limits, eventual consistency, or flaky services fail and the agent must decide whether to retry, reconcile state, resume, change tactic, escalate, or stop.