Data & AnalyticsHealth & Life SciencesAI & Agent WorkflowsReleased 8 Oct 2026
Builds a golden question-answer evaluation set for a RAG system from a domain corpus. Produces three artifacts — a calibration split, a held-out test split (never viewed during development), and an adversarial set (absent-topic, multi-hop synthesis, distractor-rich). Every Q-A row carries source-document IDs and source-text spans so downstream evaluation can separately attribute failures to retrieval vs. generation. Forces human review on LLM-drafted Q-As before they enter the golden set. Locks the final set with a dataset hash and SemVer. Use when starting RAG evaluation from scratch in a custom domain (legal, medical, internal product docs) where no public benchmark fit…