Senior AI Trainer — Scale AI (Remote) (Jan 2022 – Present)
Served as a Senior AI Trainer evaluating LLM responses and supporting RLHF preference-ranking workflows. Drove quality through daily inter-rater evaluation across instruction-following, summarization, creative writing, and code tasks. Helped translate evaluation findings into policy-relevant training data and edge-case coverage. • Wrote preference justifications for RLHF pipelines • Ranked model responses and captured edge cases for policy updates • Maintained 96%+ inter-rater agreement across 5,000+ tasks • Led weekly calibration sessions that reduced team error rates by 24% in 3 months