While I have not held a formal data labeling role, I have extensive hands-on experience evaluating AI model outputs acro
While I have not held a formal data labeling role, I have extensive hands-on experience evaluating AI model outputs across ChatGPT, Claude, Gemini, Grok, and other LLMs. This includes assessing response quality, factual accuracy, reasoning, instruction-following, safety, and consistency across a wide range of topics. I have also built and tested AI-powered applications, created prompt workflows, compared model performance, identified errors and hallucinations, and iteratively improved outputs through structured feedback. While my experience is primarily practical rather than professional data annotation, many of the skills overlap with AI training and evaluation work.