"Modality" is the technical word for a format of information: text, image, audio, video. A multimodal model handles more than one. You can show it a screenshot and ask what is wrong, hand it a photo of a whiteboard and get typed notes, or have it look at your ad creative and describe what a viewer sees first.
The frontier assistants (Claude, ChatGPT, Gemini) are all multimodal now: they accept images and often audio alongside text, and some generate images and speech in return. The distinction between "text AI" and "image AI" is dissolving into single systems that move between formats.
Why you keep hearing it
Because each new modality unlocks visibly impressive demos (point your camera at something and just talk about it), and because model announcements advertise it. "Natively multimodal" means the model was built from the start to handle multiple formats rather than having them bolted on.
What it means for you
Practical and immediate: you can stop describing things to AI and start showing them. Screenshot the competitor's landing page, photograph the packaging, upload the ad. Marketers work in visual material all day, and multimodality is what made AI able to look at that material with you.