Process multimodal image, audio, and video inputs with vision-language models and transcription pipelines.