AI training data annotator (chatbot evaluation) at Prolific
Annotated and rated conversational data for AI chatbot training and research studies focused on Egyptian Arabic dialect and dialogue quality. Provided detailed feedback on model responses and evaluated usability and relevance across different scenarios. Performed quality testing and synthesized findings into actionable recommendations to improve dialogue naturalness and cultural appropriateness. • Conversational annotation and response rating • Feedback on model responses for quality and cultural fit • Scenario-based evaluation of usability and relevance • Quality testing and reporting of actionable recommendations