Skip to content

Multimodal AI

Image captioning

Using AI to generate a written description of an image.

Example

An AI system drafts alternative text describing a photograph.

Why people use it

It makes pictures easier to describe, search or understand without seeing them.

What you'll hear

“Can it suggest a description for this photo?”

What this means for you

Check identities, actions and sensitive details before publishing a caption.

Can you control it?

Sometimes

Sometimes. Your choices depend on the tool and your access. The settings available to an everyday user may differ from those available to the people running it.

Common questions

Can a caption mention something that is not visible?
Yes. AI systems can infer or invent details beyond the image.
Does one picture have only one correct caption?
No. Useful captions can emphasize different details depending on who needs the description.
Can a caption identify a person reliably?
Not necessarily. A description can guess an identity incorrectly, even when other visible details are right.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice