AI Evaluator - Outlier AI
Developed a multi-phase LLM response evaluation pipeline to score outputs across multiple rating dimensions. Designed STEM prompt-creation workflows with source exclusion and multi-stage source verification to improve response reliability. Validated generation quality using go-to-first-answer checks via live web search. • Built an 8-phase LLM evaluation pipeline with anti-hallucination gates • Designed prompt-creation workflows with paywalled-source exclusion and verification • Ran GTFA validation using live web search for difficult cases • Applied LLM evaluation and web validation techniques to improve model outputs