Improve long-form text-to-speech through a running Voicebox REST server by planning semantic speech segments, generating them separately, preserving provenance, and merging compatible WAV audio. Use when a user explicitly asks to use Voicebox, a local cloned voice, or an authorized Voicebox profile for a long narration, article, script, voiceover, or multi-paragraph TTS asset. Do not use for generic TTS provider selection or ChatCut built-in cloud voices unless the user specifically chooses Voicebox.