Use when implementing a vision-language model — the LLaVA-style recipe of a vision encoder, an MLP projector, and a decoder LM, fused by prepending projected image tokens to text.