AI Training & Evaluation Specialist — Scale AI (2021–2024)
Evaluated large language model outputs for accuracy, factual consistency, safety compliance, and overall response quality in healthcare-focused contexts. Supported the improvement of dataset reliability through iterative annotation and refinement workflows aligned to prompt/rubric expectations. Ensured outputs met quality and safety criteria before downstream use. • Assessed LLM responses against reference facts and consistency criteria • Flagged potential safety concerns and non-compliant content • Participated in annotation pipeline contributions and workflow refinement • Worked with healthcare-related datasets to improve reasoning reliability