AI Training Specialist, Scale AI
Performed RLHF data collection by ranking model outputs and creating preference judgments aligned to helpfulness, honesty, and harmlessness. Wrote justifications for subtle quality differences to improve downstream model behavior and safety. Ensured consistency with rubrics while identifying edge cases that impacted preference outcomes. • Ranked responses for pairwise preference datasets • Authored detailed preference rationale text • Audited failures to flag ambiguous or unsafe cases • Contributed to guideline updates to reduce edge ambiguity