Multimodal AI
Computer vision
The area of AI that helps computers find and interpret information in pictures and videos.
Example
A camera system detects damaged products on a production line.
Why people use it
It lets software work with pictures and video at a scale people could not manage manually.
What you'll hear
“The camera system checks products for damage.”
What this means for you
Test lighting, viewpoints and real-world examples relevant to the task.
Can you control it?
Developer-only
The people building or running the AI choose this setup. An everyday user generally needs their help to change how this part works.
Common questions
- Does computer vision see the world like a person?
- No. It processes visual patterns and may fail under unfamiliar conditions.
- Can it understand what is outside the picture?
- It may guess from visible clues, but missing parts of a scene are not directly observed.
- Why might it fail in a new location?
- Different lighting, backgrounds or camera positions can make familiar objects look unfamiliar to the system.