I have experience working on AI training and data labeling tasks, including text annotation, response quality evaluation
I have experience working on AI training and data labeling tasks, including text annotation, response quality evaluation, and preference ranking between model outputs. My work has involved reviewing AI-generated content for accuracy, relevance, and instruction-following, as well as flagging responses that contain factual errors, unsafe behavior, or policy violations. I'm familiar with evaluating outputs across multiple dimensions — tone, helpfulness, factuality, and adherence to specific constraints — and providing structured feedback that helps improve model performance. More recently, I've been involved in agentic AI evaluation, where I assess how well models plan and execute multi-step tasks, coordinate across tools, and produce verifiable end-state artifacts. This includes identifying safety failures in model behavior, building evaluation rubrics, and rating model outputs based on outcome-focused criteria. I'm comfortable working with detailed guidelines and applying consistent judgment across a high volume of tasks.