Freelance Data Labelling Specialist – RLHF & Evaluation
I participated in RLHF (Reinforcement Learning from Human Feedback) by ranking and evaluating AI-generated responses. My work involved judging response quality, accuracy, and safety to optimize generative model outputs. This contributed to the refinement and safety of conversational AI systems. • Assessed AI output for relevance, correctness, and appropriateness. • Provided human feedback to improve model alignment. • Followed strict guidelines and safety protocols. • Completed multiple RLHF evaluation batches with high accuracy.