Design and calibrate an LLM-as-judge grader — rubric, prompt, bias controls, and validation against human labels. Use when eval outcomes can't be checked programmatically, when judge scores disagree with human judgment, or when setting up pairwise model comparisons.