Senior AI Engineer — Golden Dataset & HITL Human Evaluation, Preference Ranking (Accenture — Enterprise AI Platform)
Designed and maintained LLM evaluation datasets and CI gates, including stratified golden dataset construction with adversarial query buckets. Built HITL human evaluation workflows for clinical and compliance-sensitive outputs using LangGraph with structured reviewer actions. Delivered preference ranking and pairwise evaluation annotation programs to support prompt regression testing and A/B decisions. • Maintained a 100-query golden eval suite with explicit expected outputs, evaluation dimensions, and automated CI blocking thresholds. • Engineered HITL review queues where reviewers approve/modify/reject with confidence selection, source verification, and hallucination flags. • Conducted preference ranking annotation with calibrated anchor examples and Kappa gating to keep annotator drift under control. • Used quarterly refresh and bad-regression discovery via reviewer feedback to drive golden dataset updates and prompt iteration cycles.