Reinforcement Learning from Human Feedback (RLHF)
Learning from Human Feedback (RLHF) Reinforcement Learning from Human Feedback (RLHF) is a machine learning approach that uses human preferences and evaluations to guide an AI model toward producing more useful, safe, and desirable outputs.
What is Reinforcement Learning from Human Feedback (RLHF)?
RLHF typically involves collecting human feedback on model-generated responses, training a reward model to represent those preferences, and then optimizing the AI model using reinforcement learning. This allows the model's behavior to be adjusted based on how people evaluate its outputs.
Why is Reinforcement Learning from Human Feedback (RLHF) Important?
RLHF helps align AI models with human preferences that can be difficult to express through conventional training objectives. It can improve instruction following, response quality, helpfulness, and safety, although the resulting behavior depends on the quality and coverage of the human feedback.
Common use cases
RLHF is commonly used for training and aligning large language models, conversational AI, AI assistants, instruction-following systems, and other generative AI applications.