AI Trainer / Core Contributor — Outlier AI (Remote)
Evaluated multi-turn AI model outputs for accuracy, structural integrity, and logical flow in high-tier technical domains. Graded responses using complex project rubrics and authored clear step-by-step written justifications for rating scores. Identified hallucinations, data inconsistencies, and edge cases to improve dataset quality and training effectiveness. • Rated and reviewed LLM answers against prompt constraints • Produced rationale for grading outcomes with rigorous reasoning • Debugged generated code and mathematical workflows for compliance • Flagged subtle failure modes to guide iterative dataset improvements