AI buzzword dictionary
What is Multimodal AI?
AI that can work with more than one type of input or output, such as text, images, audio and video.
What this actually means
A multimodal assistant might read a photo of a document, listen to a question, and reply in text or speech.
It widens what you can ask, but each type of input brings its own chances for error — a blurry photo or a noisy recording still causes mistakes.
Example
You photograph a meter reading and ask, 'What is the reading and how much has it increased since 45,210?'
Where you may meet it
Phone assistants, accessibility tools that describe images, apps that translate signs through the camera.
Last reviewed: 25 September 2026
