LLM Evaluation & Annotation
Worked at Turing on LLM evaluation, RLHF annotation, SFT trajectory work, and agent function-call evaluation. Assessed model outputs for accuracy, instruction following, and response quality using structured rubrics. Delivered detailed written feedback and consistently met quality standards across multiple projects.