Design and review evaluation workflows for LLM, RAG and AI systems, including test sets, metrics, judge prompts, reliability, bias, regression testing and failure analysis.