Skip to content

Model Serving & Economics

Inference endpoint

A network address where applications can send requests to a deployed model.

Example

An application sends prompts to a stable HTTPS endpoint backed by a deployed model.

Why people use it

Knowing how “Inference endpoint” works helps teams balance model quality, reliability, speed, and cost.

What you'll hear

“How does Inference endpoint affect latency, quality, or cost at scale?”

Related terms