AI Data Annotator / LLM Evaluator (Scale AI via Outlier)
Evaluated LLM-generated responses by reviewing, analyzing, and ranking reply quality using detailed guidelines and rubrics. Provided constructive written feedback to improve helpfulness, correctness, and alignment with expected behavior. Followed project requirements to maintain high consistency and accuracy across high-volume remote tasks. • Reviewed AI model responses and chains of thought for correctness and approach • Ranked outputs based on quality dimensions (correctness/quality) per rubric • Authored feedback to improve model alignment and performance • Completed consistent high-volume evaluation work in a remote setting