Multimodal AI
Vision-language model
VLM
A model that jointly understands visual information and language.
Example
A vision-language model describes an image and answers questions about its contents.
Why people use it
Teams use “Vision-language model” when they need to choose systems that handle the required media correctly.
What you'll hear
“Does this model support Vision-language model, or is it limited to text?”