Lead AI Evaluator & Data Annotation Specialist, Scale AI (Remote)
Provided evaluation, ranking, and rating of AI-generated responses for safety, accuracy, tone, and helpfulness as part of RLHF workflows. Reviewed and annotated model outputs with precise adherence to complex guidelines, flagging subtle errors and edge cases. Produced structured written feedback to improve downstream training quality and consistency across high-volume tasks. • RLHF ranking and safety/helpfulness evaluation • Prompt-response evaluation with structured critique • Ground-truth validation and dispute resolution • Guideline refinement to reduce inter-annotator disagreement