Extract one speaker's complete, source-cited dialogue corpus from a folder of speaker-attributed sources — movie scripts, TV transcripts, stage plays, or prose novels — with countable voice metrics (turn-length distribution, short-turn share, aria frequency, profanity rate), as the deterministic front end to voice-profile-builder. For prose it can emit a separate narration corpus alongside the dialogue. Use when someone wants to pull every line a character or real person speaks across a source set in order to build or refine a voice profile, or says "extract X's corpus", "pull every line X says", "build the voice corpus for X", "get all of X's dialogue", or needs analysis…