AI Data Evaluator | Turing Inc: Deployed to Google Gemini
As an AI Data Evaluator for Turing Inc. deployed to Google Gemini, I delivered RLHF preference rankings and systematic QA across LLM outputs. My responsibilities included hallucination and error detection, classification tasks, and SFT dataset validation for high-stakes domains. I performed gold standard QA and flagged content safety violations across a variety of sensitive subject areas. • Delivered 500+ RLHF preference rankings per week using structured multi-criteria rubrics • Detected and classified hallucinations, factual errors, confabulations, and bias • Conducted QA validation to ensure annotation consistency and reliability • Ensured adherence to safe messaging and flagged harmful or biased content