Design and review agent runtime observability and model/provider fault tolerance. Use when adding metrics, traces, cost tracking, audit alignment, structured runtime events, model fallback, structured output repair, refusal handling, provider timeouts, or local execution tree diagnostics.