Maisa Korhonen
← Back to glossary

Multimodal model

An AI model that works across formats: it can see images, hear audio or watch video, not just read text.

"Modality" is the technical word for a format of information: text, image, audio, video. A multimodal model handles more than one. You can show it a screenshot and ask what is wrong, hand it a photo of a whiteboard and get typed notes, or have it look at your ad creative and describe what a viewer sees first.

The frontier assistants (Claude, ChatGPT, Gemini) are all multimodal now: they accept images and often audio alongside text, and some generate images and speech in return. The distinction between "text AI" and "image AI" is dissolving into single systems that move between formats.

Why you keep hearing it

Because each new modality unlocks visibly impressive demos (point your camera at something and just talk about it), and because model announcements advertise it. "Natively multimodal" means the model was built from the start to handle multiple formats rather than having them bolted on.

What it means for you

Practical and immediate: you can stop describing things to AI and start showing them. Screenshot the competitor's landing page, photograph the packaging, upload the ad. Marketers work in visual material all day, and multimodality is what made AI able to look at that material with you.

Updated 12 July 2026