Data Quality Analyst & Benchmark Specialist (Anuttacon, Remote)
Owned end-to-end quality assurance for 16 benchmark projects by designing test cases and answer sets from scratch across multiple content domains. Conducted comparative preference rankings and pass/fail evaluations using established rubrics to generate RLHF training data that informs model fine-tuning pipelines. Coached and calibrated 105 annotators to maintain consistent scoring quality across distributed teams. • Designed benchmark test cases and answer sets across food, media, art, and lifestyle content. • Performed rubric-based preference ranking and pass/fail evaluation for RLHF training data. • Coordinated with quality analysts to identify systemic errors and refine scoring rubrics. • Served as subject matter expert for video tasks, resolving ambiguities and refining SOPs to reduce recurring errors.