Skip to content

Multimodal AI

Voice activity detection

Detecting which parts of a recording contain someone speaking.

Example

A speech tool skips long silences before turning speech into writing.

Why people use it

It helps audio tools focus on speech instead of processing long stretches of silence.

What you'll hear

“Start processing when someone begins speaking.”

What this means for you

Try quiet voices and noisy rooms to check that speech is not cut off.

Can you control it?

Developer-only

The people building or running the AI choose this setup. An everyday user generally needs their help to change how this part works.

Common questions

Does voice activity detection identify the words being spoken?
No. It detects speech presence rather than transcribing content.
Can it mistake background sound for speech?
Yes. Music, noise or other sounds can trigger an incorrect detection.
Can it cut off a quiet speaker?
Yes. Faint speech may be mistaken for silence or background noise.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice