AI Model Evaluation Specialist
As an AI Model Evaluation Specialist at Appen, I assessed and rated outputs from large language models to improve their performance. I utilized human-in-the-loop evaluation techniques to provide detailed feedback and refinement to model responses. I delivered structured insights and evaluation reports to support continuous model improvement. • Conducted output review for accuracy, relevance, and safety • Identified and documented inconsistencies, biases, and factual errors • Applied reinforcement learning from human feedback concepts • Collaborated with Appen's proprietary tools and workflows.