AI Model Evaluator (RLHF / Quality Assurance) - DataAnnotation.tech
You evaluated and graded complex multi-turn large language model responses to ensure factual accuracy, logical soundness, and adherence to safety and stylistic guidelines. You authored high-quality ground-truth responses and developed comparative evaluations to document subtle differences in tone, bias, and correctness. You performed quality assurance by detecting hallucinations, logical fallacies, and safety violations to support iterative refinement of next-generation AI systems. • Assessed multi-turn LLM outputs for factual integrity, reasoning quality, and policy compliance. • Generated ground truth for complex instruction-following and training use cases. • Conducted side-by-side comparative analyses to guide model updates. • Flagged hallucinations, unsafe content, and other quality issues for remediation.