AI Training & Evaluation Specialist (Contract) — Mercor
Provided RLHF annotation and evaluation for frontier model training cycles, including rating and ranking LLM outputs across reasoning, coding, math, and creative writing. Achieved high-quality performance and contributed to iterative safety improvements through systematic red-teaming deliverables. Developed reusable prompt templates and evaluation rubrics to improve labeling efficiency for a team of contractors. • Completed 3,000+ RLHF annotation tasks with a 98.4% quality score • Rated and ranked LLM responses for multiple model training cycles • Built evaluation rubrics and prompt templates, reducing per-task time by 30% • Delivered red-teaming reports identifying 45+ edge cases and safety vulnerabilities