Using an LLM to compare two candidate outputs (A/B) against a rubric instead of scoring each in isolation, to reduce judge score variance.