Use when testing chatbots, voice assistants, or conversational AI agents specifically — multi-turn conversation flow, context/memory retention, intent recognition, STT/TTS accuracy, interrupt handling, and fallback behavior. Builds on llm-testing for the underlying model concerns.