Generate test scenarios and evaluate test call transcripts for voice agents. Use after deployment to assess real-world performance. Feeds findings back into prompt iteration. Requires human test callers.