Multimodal AI
Native multimodality
AI designed to learn from and work with several kinds of material, such as words, pictures and sound.
Example
An AI system learns relationships between text and audio in a shared system.
Why people use it
It lets one AI system work across different kinds of information more directly.
What you'll hear
“Can it connect what it hears with what it sees?”
What this means for you
Test the specific combinations of media your workflow uses.
Can you control it?
Sometimes
Sometimes. Your choices depend on the tool and your access. The settings available to an everyday user may differ from those available to the people running it.
Common questions
- Does native multimodality imply equal strength in every modality?
- No. Capabilities and limitations can differ greatly across tasks.
- Does it have to support every media type?
- No. A system may directly handle some combinations, such as text and images, without supporting others.
- Can combining media introduce new mistakes?
- Yes. It may connect a sound, picture or piece of text to the wrong thing.