Use when designing or auditing a PLDI evaluation — choosing defensible benchmark suites and baseline compiler configurations, measuring runtime, compile time, and memory with warmup and variance discipline, running ablations that isolate the claimed mechanism, and scoping claims to the platforms measured.