Senior AI Data Annotation Specialist (RLHF pairs, instruction-following, red-teaming evaluations)
Delivered 18,000+ RLHF comparison pairs for a top-3 LLM provider, driving a 97.4% inter-rater agreement score over a 6-month engagement. Led multilingual instruction-following dataset annotation for English, German, and Edo, establishing guidelines used by the wider project team. Performed red-teaming evaluations focused on safety and hallucination detection to inform model policy updates. • Created and validated RLHF pair labels (comparison-style human feedback) for large-scale language modeling. • Conducted code generation quality assessments by rating correctness, efficiency, and style across thousands of model outputs. • Logged and investigated edge cases (340+ flagged) to improve dataset coverage and safety. • Mentored and coordinated a team of 8 annotators on labeling consistency and rubric adherence.