Evaluation of AI outputs: hallucination detection, answer relevance, context recall, LLM-as-a-judge benchmarking, RAGAS metrics, and regression testing.