Harden an Agent Skill from its automatic agent evals: score every eval run against a fixed fidelity rubric read from the agent's own log, fix the skill at the root cause of each deviation, cut a patch release, let the release re-run the evals, and repeat until every agent scores full.