Send and receive multimodal content in AIVAX chat completions — images, audio, video, files — including when to use direct media vs. multimodal preprocessing to text. Load when the user input or the model output includes media, or when the model must reason over attached content.