AI Evaluation Generalist (Freelance) – Data Annotation (AI/LLM evaluation)
Freelance AI Evaluation Generalist performing quality and safety assessments of large language model (LLM) outputs. Labeled and compared model responses for helpfulness, reasoning quality, factual accuracy, clarity, and instruction adherence. Identified inaccuracies and hallucinations using detailed evaluation guidelines. • Evaluated response helpfulness and overall quality • Assessed reasoning and factual accuracy • Checked for hallucinations and instruction following • Followed written evaluation protocols across varied projects