Freelance AI Trainer & LLM Response Evaluator (Multiple Annotation Platforms)
Evaluated, ranked, and rewrote responses from frontier LLMs across reasoning, math, coding, creative writing, and instruction-following tasks using structured rubrics. Produced gold/ideal responses used as supervised fine-tuning data and maintained consistent top-tier quality scores on internal QA audits. Designed adversarial and edge-case prompts to expose hallucinations, refusal failures, and reasoning gaps for ongoing red-team batches. • Performed pairwise preference tasks and multi-turn dialogue evaluations • Managed inter-rater agreement above platform thresholds across multi-week projects • Provided written feedback to refine rubric guidelines • Focused on factuality, citation quality, and rubric-based QA scoring