Skip to content

Multimodal AI

Text-to-audio

Generating audio from a written prompt or description.

Example

A tool creates the sound of rainfall from a written request.

Why people use it

It helps creators produce sounds from a description instead of recording them directly.

What you'll hear

“Create the sound of rain on a roof.”

What this means for you

Listen to the result and check the permitted uses before sharing it.

Can you control it?

Sometimes

Sometimes. Your choices depend on the tool and your access. The settings available to an everyday user may differ from those available to the people running it.

Common questions

Is text-to-audio only text-to-speech?
No. It can include music, sound effects and other audio.
Can the same description produce different sounds?
Yes. Different attempts or settings can change the generated result.
Does it recreate a specific real recording?
Not necessarily. It creates a sound matching learned patterns and the request, rather than verifying a real event.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice