Reinforcement Learning with Human Feedback (RLHF) by Hand ✍️
Calculating AI by Hand: 28 of 28
Library › Calculating AI by Hand ✍️
Reinforcement Learning with Human Feedback (RLHF) by Hand ✍️
Reinforcement Learning from Human Feedback (RLHF) is a popular technique to ensure that an LLM aligns with ethical standards and reflects the nuances of human judgment and values.
Without RLHF, an LLM relies only on data and would think doctors must be men, because the data likely reflects existing biases in our society.
With RLHF, an LLM is given human feedback that doctors can be both man and women. The LLM can update its weights until it begins to use "them" rather than "him" to refer to a doctor.
Moreover, we hope the LLM not only addresses the specific bias about doctors but also learns the underlying value of "gender neutrality" and applies it to other professions, for example, learns to use "them" to refer to a CEO, even though it wasn't explicitly taught by a human.
Claude 3 released by Anthropic, sets a new high bar for safety standards. It uses an advanced technique called "Constitutional AI" by extending RLHF, enhancing the H in RLHF with AI.
How does RLHF work?
Setup
Step 1 of 15: Given
Reward Model (RM)
Large Language Model (LLM)
Two (Prompt, Next) Pairs
Train Reward Model
Goal: Learn to give higher rewards to winners
Step 2 of 15: Preferences
A human reviews the two pairs and picks a "winner"
(doc is, him) < (doc is, them) because the former has gender bias.
Steps 3-6: Calculate the Reward for Pair 1 (Loser)
Step 3 of 15: Word Embeddings
Lookup word embeddings as inputs to the RM





