Audit benchmark, model, harness, product, or performance claims before release or publication. Use when results include scores, pass rates, tokens, time, cost, confidence intervals, censored runs, internal tasks, or cross-provider comparisons and the claim must not exceed the evidence.