Create new skills, modify and improve existing skills, and measure ONE skill's performance. Use to: create a skill from scratch, edit or optimize an existing skill's description for triggering accuracy, draft test prompts for a new skill and run models against them, measure variance of a single skill (same prompt N runs — how consistent is it?), or edit a skill's evals/trigger-eval.json (e.g. add near-miss negatives) and rerun that skill's eval. Single-skill authoring/eval/variance work is THIS skill; whole-library benchmarks against real usage history are the benchmark skill.