AI Model Training Specialist (Contract) - Mercor AI
Evaluated and ranked LLM outputs for reasoning quality, factual accuracy, coherence, and domain alignment across technical and research-focused tasks. Drove RLHF workflows by producing structured evaluation feedback and identifying complex edge-case failure modes. Performed QA on structured and multimodal datasets while maintaining strict guideline compliance and improving response quality through systematic prompt optimization and adversarial testing. • LLM output evaluation across safety, accuracy, and reasoning dimensions • RLHF workflow support with structured feedback generation • Prompt optimization, adversarial red-teaming, and hallucination detection • Multimodal (text, audio, video) dataset annotation QA and benchmarking