AI Model Evaluation Specialist
As an AI Model Evaluation Specialist at Appen, I assessed and rated outputs from large language models to improve response accuracy, relevance, and safety. I provided detailed feedback to refine model behavior using human-in-the-loop evaluation techniques. I contributed to iterative model improvement by delivering structured evaluation reports and insights. • Evaluated model-generated responses for quality and compliance • Identified gaps such as bias and factual inaccuracy • Applied RLHF methodologies to drive improvements • Enhanced AI system reliability through structured insights