Skip to content

Multimodal AI

Vision-language model

VLM

A model that jointly understands visual information and language.

Example

A vision-language model describes an image and answers questions about its contents.

Why people use it

Teams use “Vision-language model” when they need to choose systems that handle the required media correctly.

What you'll hear

“Does this model support Vision-language model, or is it limited to text?”

Related terms