Use when the user asks how to do error analysis on agent traces — "my agent fails in production and I don't know why," "how many traces should I read," "how do I cluster agent failures," "which failures should I fix first," "should an LLM triage my traces," or "how do I turn production failures into tests." Walks the recurring trace-reading cycle: reading full traces with a saturation stopping rule, clustering failures into a taxonomy (derive vs import), attributing clusters to environment vs grader vs model, prioritizing fixes by frequency and impact, and promoting one representative case per cluster into regression tests. Not for localizing a specific live regression to…