Skip to content

Model Training & Adaptation

Direct preference optimization

DPO

A technique that trains a model directly from preferred and rejected response pairs.

Example

Training directly increases the probability of preferred responses over rejected alternatives.

Why people use it

Teams use “Direct preference optimization” when they need to adapt models deliberately and diagnose training problems.

What you'll hear

“Would Direct preference optimization improve the model for our specific use case?”

Related terms