AI Trainer — Mercor
Provided expert evaluations and ratings for LLM outputs to improve accuracy, reasoning quality, and code correctness. Performed RLHF-related evaluation tasks including comparative response assessment. Identified failure patterns such as hallucinations and edge cases to guide iterative model refinement. • Accuracy and reasoning quality scoring • Code correctness checks for generated solutions • Comparative response rating for RLHF • Reporting of hallucinations and reasoning failures