Multimodal AI: Working With Text, Images, Audio, and Video
June 19, 2026 · 8 min read
Multimodality removed the transcription step between the real world and the model.
Vision
Screenshots, whiteboards, and receipts become directly usable inputs.
Audio
Voice conversation removes typing friction and enables hands-free workflows.
Video understanding
Long-video question answering is newly practical for lectures and recordings.
Practical tip
Combining modalities in one prompt — image plus instructions — beats describing the image in words.