LLM response evaluation & factuality
Small contract piece: judge prompt + two AI responses. Main tasks were rubric scoring, checking whether answers matched facts and details (lots of lightweight research/verification), and following the evaluator guidelines consistently. Scope was modest—repeatable workflows, unpredictable content. Often, one answer was nearly there while the other had obvious factual or logic problems—so diligence mattered.