LLM Evaluator / Agent Output Rater
I evaluated and rated LLM and multi-agent pipeline outputs for relevance, grounding, and safety. I compared outputs against ground-truth legal standards and user health information. I iterated on prompt and agent instructions to address systematic model errors found in evaluations. • Scored LLM outputs for accuracy and appropriateness using structured rubrics • Conducted bias, hallucination, and failure mode analyses for each agent output • Documented evaluation findings for product and research teams • Optimized scoring guidelines and uncertainty calibration for legal eligibility and health domains