Skip to content

Model Serving & Economics

Cost per request

The average expense incurred each time an application calls an AI model or service.

Example

A team divides its monthly model bill by completed requests to estimate cost per request.

Why people use it

Understanding “Cost per request” helps teams balance model quality, reliability, speed, and cost.

What you'll hear

“How does Cost per request affect latency, quality, or cost at scale?”

Related terms