AI Model Trainer & Evaluator | Outlier AI (Scale AI)
Trained and fine-tuned large language models using structured RLHF workflows, producing gold-standard instruction data to improve performance on reasoning, coding, and creative tasks. Authored thousands of high-quality prompt-response pairs and iteratively refined training examples based on evaluation outcomes. Performed comparative response assessments and ranked outputs using HHH (helpfulness, harmlessness, and honesty) criteria. • Created prompt-response datasets for instruction-tuned models (SFT/RLHF preparation) • Evaluated outputs via side-by-side helpfulness/harmlessness/honesty scoring and ranking • Reviewed and rated AI-generated code submissions for correctness and best practices • Participated in domain-specific training sprints (mathematics, logical reasoning, STEM) to expand coverage