AI Evaluation Contributor - Mercor
Completed AI evaluation work across Project Neon and Project Nova, including RLHF annotation, SFT review, generalist model evaluation, multimodal analysis, task flagging, and rubric-based response assessment. Evaluated AI-generated outputs for helpfulness, instruction following, factual consistency, reasoning quality, safety, alignment, clarity, completeness, and user-intent fulfillment. Also evaluated audio and visual annotation tasks by applying detailed guidelines to assess multimodal content quality, artifacts, alignment, and task-specific labeling requirements. • Performed comparative RLHF judgments across candidate responses and identified hallucinations, logic gaps, and rubric deviations • Reviewed SFT-oriented outputs for ideal-response quality, structure, tone, and requirement alignment • Applied safety-first, policy-aware judgment when flagging problematic or non-compliant content • Ensured consistent long-session performance through precise guideline interpretation