Use when validating whether an AI agent skill, workflow, prompt, toolset, or autonomous coding loop actually improves results across repeatable tasks.