Senior Data Labeler & AI Trainer (LLM safety evaluation)
Evaluated 10,000+ model responses for helpfulness, harmlessness, and honesty to support LLM safety improvements. Applied structured criteria to identify quality and safety issues in generated text. Contributed to reinforcement learning from human feedback prompt evaluation activities. • Rated responses using safety and helpfulness dimensions. • Identified failure modes affecting honesty and safety. • Supported RLHF prompt evaluation workflows. • Helped teams refine datasets for safer LLM outputs.