AI Evaluation Task Writer (LLM evaluation / golden solutions / scoring rubrics) | Mercor
Authored complex, domain-specific LLM evaluation tasks to test model performance on policy/claims domain scenarios. Built evaluation packages including prompts, weighted scoring rubrics with critical-value gates, and verified golden answer keys. Executed and graded model runs across multiple engines while logging failures by severity to quantify performance and calibrate task difficulty. • Created prompt suites and reference memos reflecting regulatory compliance and claims authority limits • Developed Trajectory scoring and critical-error analysis to assign competency bands • Used staging/integration-style validation concepts to ensure evaluation correctness before approval • Refined tasks through a peer review and quality-control pipeline