Freelance AI Evaluation & Data Annotation (Remote)
Evaluated AI-generated responses by rating accuracy, relevance, and completeness against provided instructions and rubrics. Compared multiple model outputs to identify hallucinations, bias, and logical inconsistencies before selecting best-performing responses. Produced clear written feedback intended to improve downstream model performance and consistency. • Applied structured scoring rubrics and annotation standards • Performed comparative evaluation across multiple outputs • Detected and flagged hallucinations and bias • Delivered written improvement notes for model refinement