Prompt Evaluator, Outlier Remote (Jan 2023 - Present)
Evaluated AI-generated responses to complex medical and scientific prompts, assessing clinical accuracy, reasoning quality, and alignment with evidence-based standards. Developed and reviewed challenging domain-specific questions across pharmacology, anatomy, infectious disease, and public health. Tagged and categorized clinical content by specialty and subcategory to support routing for quality review, while providing structured feedback to guide iterative model improvements. • Assessed clinical accuracy and reasoning quality of model outputs • Authored/reviewed domain-specific medical and scientific questions • Tagged and categorized content for expert review routing • Wrote structured feedback identifying factual and reasoning gaps