Model Training & Adaptation
Reinforcement learning from human feedback
RLHF
Training that uses human preference judgments to shape model behavior.
Example
People rank two answers, and those preferences help train the assistant's behavior.
Why people use it
Knowing how “Reinforcement learning from human feedback” works helps teams adapt models deliberately and diagnose training problems.
What you'll hear
“Would Reinforcement learning from human feedback improve the model for our specific use case?”