Aiodic

Back to the table

38 Intermediate Rh RLHF Data & Training

RLHF

Reinforcement learning from human feedback.

People rate model responses, and those preferences train the model to be more helpful and less harmful. A key step in making chat assistants.