AI Training Data Specialist | Outlier AI
Performed text classification and response ranking to create RLHF training data for large language models and AI systems. Annotated AI-generated outputs for correctness, tone, safety, and instruction-following accuracy using detailed rubrics across large volumes. Ensured dataset integrity by identifying edge cases, ambiguous prompts, and low-quality outputs while meeting strict turnaround timelines. • Generated training labels via classification and response ranking tasks for RLHF workflows • Evaluated outputs on multiple dimensions including correctness, safety, and intent adherence • Applied rubric-based guidance to maintain high inter-annotator agreement • Flagged edge cases and inconsistencies to protect training data reliability