Skip to content

Infrastructure & Compute

Time to first token

The wait before an AI starts sending the first small piece of its reply.

Example

A user waits two seconds before an answer begins appearing.

Why people use it

It measures the wait before an AI answer starts appearing.

What you'll hear

“Why is the screen blank for so long?”

What this means for you

Compare both initial waiting time and time to a complete answer.

Can you control it?

No

No direct control. This describes a wider issue, concept or result rather than something you can simply switch on or off in a tool.

Common questions

Is time to first token the total response time?
No. Generating the rest of the answer takes additional time.
Can a long question increase the initial wait?
Yes. The system may need more time to read and process a long request.
Can two services start equally fast but finish differently?
Yes. One may produce the remaining answer much faster after the first word appears.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice