Use when testing and evaluating AI agents and tool-using LLM systems — task-success and trajectory evaluation, tool-call correctness, LLM-as-judge calibration, regression datasets, non-determinism (pass@k), cost/latency budgets, and CI evals with DeepEval, promptfoo, LangSmith, Inspect or MLflow.