Audits a claim about model or system performance against six integrity domains — test-set contamination, baseline pinning, seed and run variance, evaluator independence, metric selection, and whether the test set is large enough for the reported gap — and emits a 0–5 reliability score with named risk tags and a replication test. Use when a vendor post, paper or release note asserts a benchmark number: "91% on GSM8K", "outperforms GPT-4o", "SOTA on MMLU", "audit this benchmark claim". Not for grading a clinical or social-science trial (use `assess-study-bias`) and not for qualitative claims such as "better reasoning".