Mercor | Data Scientist - AI/ML Expert, Contract | Remote, USA / Scottsdale, AZ
Designed frontier model evaluation workflows using rubric-based grading and reference solutions to systematically expose hallucinations, brittle code, and missed edge cases. Created benchmark-quality prompts and correction rationales to enable reproducible technical scoring across diverse test conditions. Led stress testing of AI outputs by iterating through data cleaning, feature engineering, experiment design, and production-style debugging workflows. • Built grading rubrics and reference solutions for evaluation • Developed evaluation prompts and golden sets • Performed failure analysis for edge cases and reasoning gaps • Used Python/SQL/ML tasks to probe model assumptions