Multimodal Models
Images, audio, and video join the token stream: what multimodal actually means inside the model, and what it changes at the meter.
Content last verified 2026-09.
Lessons
Sources
- Qwen2.5-Omni-7B model card — modalities in and out, as documented
- Llama 3.2 11B Vision Instruct model card — the input/output modality table
- Gemma 3 4B model card — image normalization and per-image token count
- Phi-4-reasoning-vision-15B model card — documented vision encoder and fusion
- Mitchell et al. (2018) — Model Cards for Model Reporting