Audits AI models, zero-token reflex engines, and evaluation pipelines to detect and prevent benchmark test cheating, test-set memorization, hardcoded question overrides, and artificial logit manipulation. Use this skill whenever evaluating benchmark claims, auditing suspect evaluation numbers that seem "too good to be true", designing new benchmark suites, adding anti-overfitting CI guards, or refactoring inference engines to ensure genuine generalizability.