Skip to content

Model Training & Adaptation

Reinforcement learning from AI feedback

RLAIF

Training that uses AI-generated preference or critique signals instead of only human feedback.

Example

A stronger model critiques candidate answers that are then used to improve another model.

Why people use it

Understanding “Reinforcement learning from AI feedback” helps teams adapt models deliberately and diagnose training problems.

What you'll hear

“Would Reinforcement learning from AI feedback improve the model for our specific use case?”

Related terms