Review only runs marked "Successful" to hunt for instances where the agent hallucinated a success message but quietly failed to execute the actual task. Use when reviewing an eval harness that has high scores but suspicious downstream bugs, or when an agent claims to have completed a task but state hasn't changed.