Data Annotation Specialist – Labelbox (Contract / Remote)
As part of my responsibilities, I evaluated and ranked AI-generated text responses using Reinforcement Learning from Human Feedback (RLHF) and Supervised Fine-Tuning (SFT) practices. The work focused on improving model reasoning, factuality, and alignment by applying detailed evaluation protocols. My contributions supported iterative optimization of large language model performance. • Reviewed and rated text responses for relevance and accuracy. • Applied RLHF and SFT methodologies to enhance LLM outputs. • Provided structured feedback to calibrate prompt and response guidelines. • Collaborated with cross-functional remote teams for continuous process improvement.