Guards autonomous research loops against the failure modes that produce confident wrong conclusions: reward hacking, fabricated numbers when code fails, overfitting the agent's own search signal, data leakage, undertuned baselines, single-seed claims, LLM-judge bias, and premature completion.