A/B-test a Claude Code skill against an unstructured baseline to measure real token savings. Per skill, runs two Sonnet sub-agents in fresh sandboxed worktrees — arm A solves the task cold, arm B follows SKILL.md — and compares `usage.total_tokens`. Measures one skill, several, or a whole repo in one parallel batch; emits a JSON measurement, a markdown report, and an A4 PDF. Use for "calibrate this skill", "measure tokens saved", "A/B-test the skill", "how much does X actually save", or before any public savings claim. Replaces editorial guesses with real numbers.