Diagnoses LLM failures by capturing full traces, classifying failure modes, bisecting prompt and model changes, and inspecting token and tool-call state. Use when output quality drops unexpectedly, when agents loop or stall, or when a regression appears after a model or prompt change.