AI Evaluation & Research Specialist — Infosys
Provided enterprise-scale evaluation of large language model outputs by performing structured human-feedback-style quality checks across multiple risk and correctness dimensions. Labeled and rated model responses for reasoning quality, factuality, hallucination presence, and contextual alignment to support RLHF-style preference ranking. Built and validated evaluation rubrics, reviewer calibration standards, and annotation QA processes to ensure consistent dataset quality. • Reasoning assessment and hallucination detection labeling • Factuality verification and contextual alignment review • RLHF-style preference ranking workflow contribution • Multilingual evaluation and taxonomy/synthetic data review