Use when evaluating AI-agent behavior, executable targets, reproducible regressions, sandbox readiness, or canonical sandbox evidence with AgentCI.