LLM Training Data Evaluation and Expert Adjudication
Conduct expert evaluation, review, and adjudication of training data used to improve the performance and alignment of large language models (LLMs). Assess both AI-generated and human-authored content for factual accuracy, reasoning quality, self-containment, rubric compliance, and adherence to project specifications across graduate-level benchmark tasks. Apply domain expertise in Anthropology, Sociology, Linguistics, Iberian and Latin American Studies, and related disciplines to identify subtle reasoning errors, detect hallucinations, evaluate distractor quality, and ensure that assessment items measure analytical thinking rather than simple recall or pattern matching. Provide detailed feedback to resolve reviewer disagreements, maintain annotation consistency, and support the creation of high-quality datasets for model evaluation, alignment, and instruction tuning. Applied extensive experience in higher education assessment design and rubric development to guarantee rigorous quality standards in large-scale AI training initiatives.