Use when the user asks how to build the environment their agent evals run in — "how do I isolate eval trials?", "my eval runs contaminate each other", "how do I run trials in parallel?", "how do I design realistic eval tasks?", "should the eval use the same tools/harness as production?" — or when standing up sandboxes, stateful task suites, or simulated users for agent evaluation. Walks three fidelity properties: per-trial isolation (clean, disposable, replayable sandboxes, no cross-contamination), realistic multi-tool stateful task design graded on end state, and deployed-harness identity (same tools and config as production, held constant across comparisons). Not for th…