Skip to content

Multimodal AI

Native multimodality

AI designed to learn from and work with several kinds of material, such as words, pictures and sound.

Example

An AI system learns relationships between text and audio in a shared system.

Why people use it

It lets one AI system work across different kinds of information more directly.

What you'll hear

“Can it connect what it hears with what it sees?”

What this means for you

Test the specific combinations of media your workflow uses.

Can you control it?

Sometimes

Sometimes. Your choices depend on the tool and your access. The settings available to an everyday user may differ from those available to the people running it.

Common questions

Does native multimodality imply equal strength in every modality?
No. Capabilities and limitations can differ greatly across tasks.
Does it have to support every media type?
No. A system may directly handle some combinations, such as text and images, without supporting others.
Can combining media introduce new mistakes?
Yes. It may connect a sound, picture or piece of text to the wrong thing.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice