How the evaluating side of the agentic-coding loop (the reviewer agent, the verifier subagent, the contract reviewer) grades without the leniency LLMs show toward LLM output. Covers the generator/evaluator separation, exercising the running application instead of reading the diff, rubric criteria with hard thresholds, the failure modes of talking yourself out of a finding and superficial testing, calibration examples, and the log-driven tuning loop. Use whenever an agent judges work it did not produce.