Ai Annotator, outlier (AI annotation and LLM evaluation)
Worked as an AI Annotator by designing and iterating base prompts across domains such as science, policy, and technology to drive clearer and more logically aligned responses. Evaluated AI-generated text outputs using quality rubrics to flag hallucinations, inaccuracies, and logical inconsistencies, improving reliability of prompt-response pairs. Built structured datasets of question answering, summaries, and synthesis tasks across different lengths, styles, and system-prompt configurations. • Prompt-response iteration and guideline adaptation for multiple content formats • Rubric-based evaluation including hallucination detection and factual verification • Red-teaming style failure identification and benchmarking pipeline contributions • Produced high-volume batches with consistent quality standards and asynchronous deadlines