Data Annotator (AI response evaluation and rating) — Outlier, Remote
Evaluated and graded AI-generated responses across diverse topics by checking factual accuracy, logical consistency, safety, and alignment with prompt constraints. Authored structured feedback to support iterative improvements while strictly following evolving project guidelines. Performed detailed hallucination/edge-case detection and fact-checking using authoritative sources to ensure reliability of annotated outcomes. • Graded response quality against safety and prompt-compliance requirements • Identified subtle errors, hallucinations, and edge cases in LLM outputs • Wrote clear, structured feedback to guide model iteration • Conducted meticulous fact-checking with multiple authoritative references