Designs repeatable evals for agent behavior, tool choice, trajectories, handoffs, guardrails, and final outcomes. Use when creating regression gates, benchmark datasets, trace graders, or launch criteria. Not for production telemetry alone or one-off adversarial review.