AI/Data Annotation & Evaluation Specialist — Freelance / Remote
Evaluated AI-generated responses for correctness, coherence, and contextual relevance in simulated user scenarios. Compared multiple model outputs and produced structured feedback to guide improvements in personalization and usefulness. Assessed failures related to context retention, bias, and factual grounding across multi-language outputs. • Tested personalization via prompts reflecting real-world user intents • Rated and analyzed response quality across accuracy and grounding • Documented weaknesses such as bias and context drift • Reviewed English/Kiswahili response performance for alignment