Builds and runs LLM eval suites with graded rubrics, pairwise model-graded comparison, deterministic graders, and statistical significance checks. Use when changing a model, prompt, or retrieval setup and needing evidence it helped, or when quality has regressed in production.