Test a HackerRank Orchestrate agent against its failures and inconsistencies, not just its successes — deliberately inspecting where similar cases get different treatment. Use before submission when the only testing done so far was "run it and see if the output.csv looks reasonable," when comparing how the agent handled two superficially similar tickets/claims, or when preparing concrete edge-case examples to discuss in the AI judge interview.