RLHF Reward Model
A scoring model trained on human feedback to evaluate language model generation quality.
The RLHF Reward Model is trained using comparison datasets of human preferences. It outputs scalar scores representing quality, guiding the policy modelβs alignment via reinforcement learning.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.