Extract ready avatar reference frames from talking-head videos — single face, face-region sharpness, no burned-in subtitles. Optionally saves inpaint candidates (1 face, sharp, with subtitles) in with_subtitles/ when requested or when no ready frames exist. Uses MediaPipe and EasyOCR. Use when the user asks to extract avatar frames, clean frames from video, reference images for virtual avatar, or talking-head frame extraction.