AI Evaluator (Freelance)
As a freelance AI Evaluator at DataAnnotation, I assessed outputs from large language models (LLMs) using structured quality rubrics. I wrote and refined prompts to probe and analyze model behavior, providing detailed written critiques to support model alignment. My work ensured consistent, high-quality task output under strict guidelines and tight feedback loops. • Evaluated model responses for accuracy, reasoning, instruction following, safety, and tone. • Developed prompts targeting STEM, health, general knowledge, and writing domains. • Delivered structured feedback to aid LLM training and downstream improvements. • Maintained high task volume and adherence to detailed guidelines.