Skip to content

Multimodal AI

Audio classification

Sorting recordings into groups based on the sounds they contain.

Example

AI sorts a recording as music, speech or background noise.

Why people use it

It helps sort recordings without someone listening to every file.

What you'll hear

“Which clips contain speech rather than music?”

What this means for you

Try recordings similar to the sounds the tool will hear in practice.

Can you control it?

Sometimes

Sometimes. Your choices depend on the tool and your access. The settings available to an everyday user may differ from those available to the people running it.

Common questions

Does audio classification produce a written record of speech?
Not necessarily. It predicts categories rather than the words spoken.
Can one recording receive several labels?
Yes. A clip can contain speech, music and traffic at the same time.
Will it recognize sounds it was never taught?
It may miss them or force them into a familiar category.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice