Large Language Model (LLM) Response Evaluation & Alignment
Worked as Senior AI Data Annotation Specialist with focus on the rigorous evaluation and alignment of Large Language Models (LLMs). Reviewed 50,000+ AI generated outputs with detailed, multi-dimensional quality taxonomies and complex project guidelines. Key responsibilities and achievements: - Evaluated model outputs for factuality, adherence to prompt instructions, context awareness, safety and fluency of language. - Identified, documented and categorized subtle logical fallacies, edge-case bugs, and hallucinatory behavior in model outputs to improve training loops. - Performed extensive online research and fact-verification from authoritative sources to audit the veracity of the generated content. Consistently achieved greater than 98% annotation and evaluation accuracy and regularly exceeded platform productivity and throughput benchmarks.