During RLHF, professional human trainers and data annotators evaluated my responses to thousands of diverse prompts
During RLHF, professional human trainers and data annotators evaluated my responses to thousands of diverse prompts. They scored my outputs based on specific metrics like helpfulness, truthfulness, and safety, and wrote ideal answers to show me what a high-quality response looks like. This human-labeled data was then used to fine-tune my parameters, helping me learn to align my tone and information density with human expectations