Transcribe a monolingual American English interview recording with word-level timestamps and speaker diarization, then guide the user through picking which speaker (typically the participant) to feed into alignment and vowel extraction. Produces a Rev-style JSON transcript containing only the chosen speaker. Uses faster-whisper for ASR (CTranslate2 runtime, cross-platform), SpeechBrain ECAPA-TDNN embeddings for speaker identity, and agglomerative clustering for diarization. Use whenever the user asks to transcribe an English interview, diarize a recording, or produce a Rev-style JSON transcript.