Evaluate and improve DeckForge Agent Skills by comparing baseline, current, and candidate conditions on deterministic outcome tests, trigger precision, blind review, context cost, runtime, and cross-agent consistency. Use for skill development and release gates; do not use for end-user deck creation.