AI Research Evaluator (Non-Annotation Support) - Outlier AI
Provided high-accuracy data labeling and evaluation support for large language models and generative AI systems within human-in-the-loop workflows. Conducted prompt evaluation and instruction ranking to support model performance and alignment with human reasoning. Applied strong linguistic judgment, attention to detail, and research logic to ensure quality across complex NLP tasks. • Completed summarization ranking, intent categorization, and multi-turn dialogue assessment • Delivered structured feedback to improve fluency, factuality, and ethical safety • Supported behavioral research experiments requiring precise reasoning and careful data handling • Collaborated across multiple AI evaluation platforms to meet task requirements