AI Model Evaluator (Freelance) — Outlier / Aether
Performed live Speech-to-Speech (S2S) Elo evaluations by comparing voice AI model outputs side-by-side to determine relative quality. Annotated AI-generated responses using detailed rubrics covering Helpfulness, Factuality, Instruction Following, and Style & Format. Maintained consistent evaluation standards while handling high-volume annotation with quality and productivity tracking. • Task type included relative ranking (Elo) and rubric-based scoring. • Used structured dimensions to assess response quality across multiple criteria. • Worked on high-volume evaluation batches with ongoing QC checks. • Bridged academic knowledge with practical, industry-style model assessment.