Senior AI Trainer & Language Specialist at Scale AI
Led large-scale AI evaluation and human feedback programs to improve large language model quality and efficiency across multiple domains. Evaluated tens of thousands of AI-generated responses using structured rubrics for reasoning, writing, coding, factuality, and instruction-following. Performed safety and bias verification through red-team testing and hallucination pattern analysis. • Evaluated 25,000+ AI-generated responses and provided consistent reviewer scoring using rubric-driven guidance • Developed evaluation benchmark datasets to support model alignment testing • Increased reviewer consistency by 34% through rubric design and calibration • Trained and mentored 100+ freelance evaluators for annotation/evaluation governance