Builds self-verifying experiment laboratories and measurement feedback loops for optimization, debugging, visual fidelity, quality, cost, and solution-search work. Use when a user asks to make something faster, cheaper, more accurate, less buggy, closer to a reference, or otherwise better across competing approaches where benchmarks, screenshots, traces, fixtures, or repeated trials can measure progress.