Model Training & Adaptation
Direct preference optimization
DPO
A technique that trains a model directly from preferred and rejected response pairs.
Example
Training directly increases the probability of preferred responses over rejected alternatives.
Why people use it
Teams use “Direct preference optimization” when they need to adapt models deliberately and diagnose training problems.
What you'll hear
“Would Direct preference optimization improve the model for our specific use case?”