AI Content Evaluator (Generalist) | Outlier.ai / Scale AI
Evaluated conversational model outputs for RLHF objectives, focusing on conversational flow, factual accuracy, and creative writing quality. Provided structured chain-of-thought justifications to explain nuanced differences between responses for fine-tuning. Identified and flagged subtle hallucinations and logic errors in complex multi-turn prompts while enforcing safety and ethical compliance.• Ranked model responses based on guideline criteria for RLHF ranking tasks.• Produced reasoning and justification content to support model alignment and improvement.• Audited adherence to large, evolving guideline documents (50+ pages).• Collaborated with teams on red-teaming to validate safety and compliance requirements.