Software EngineeringAI & Agent WorkflowsReleased 6 Oct 2026
Use this skill when measuring whether a Stackbone agent or workflow actually works, and whether a change made it better or worse: turning a Playground exchange into a saved eval case, building a case list and versioning it, writing a suite (its target, its criteria, its repeats, its pass mark), choosing between a deterministic scorer and an LLM judge, wiring a custom workflow as the judge, enabling the persona simulator so a case that asks the user back still gets measured, judging behaviour rather than wording with the trajectory criteria (which tools, which workflow steps, which subagents), launching a run from the CLI as a CI gate, launching a two-variant comparison fr…