CEO & Founder, Rater-X (Human-in-the-Loop evaluation infrastructure)
Built a human-in-the-loop AI evaluation framework at scale for trusted AI systems. Oversaw evaluator calibration and inter-rater agreement to ensure consistent, objective judgments across distributed teams. Translated complex evaluation guidelines into actionable workflows for multilingual and multimodal assessments. • Designed evaluation rubrics and quality frameworks for text and multimodal tasks, including naturalness scoring and dialect classification. • Recruited and trained evaluators while establishing calibration systems to maintain consistency. • Coordinated cross-functional teams to deliver reliable evaluation pipelines. • Partnered with AI companies and data organizations to improve model training data quality.