AI Trainer / Scientific AI Evaluator (Freelance) | Outlier
As an AI Trainer and Scientific AI Evaluator (Freelance) at Outlier, I evaluated and ranked AI-generated STEM responses for reasoning quality, factual accuracy, coherence, and adherence to guidelines. I produced structured written justifications and identified hallucinations, inconsistencies, and linguistic errors in model outputs. My role also included conducting multimodal evaluation tasks involving both text and image interpretation. • Generated prompts and edge cases to test model robustness • Assessed AI-generated STEM content and performed detailed response analysis • Identified and documented logical errors and hallucinations in outputs • Worked with prompt evaluation frameworks and internal tools to perform evaluations.