Score a system prompt, agent prompt, or task prompt against a stated rubric and return PASS, WEAK, or FAIL with per-dimension marks, quoted evidence, and the two or three fixes that raise the score most. Give it two versions and it reports what improved and what regressed, even when the revision wins overall. Use before shipping a prompt change, or to compare prompt versions.