AI Trainer & LLM Evaluator at Outlier (Scale AI)
Provided LLM response evaluation by ranking outputs for accuracy, tone, coherence, logical consistency, and guideline adherence. Produced prompt variants to expose weaknesses in reasoning, factuality, and instruction-following for training data improvement. Participated in RLHF by writing nuanced preference comparisons and supporting red-teaming for unsafe, biased, or policy-violating responses. • LLM response ranking and written rationales for model teams • Quality scoring against task guidelines across 500+ tasks • RLHF preference comparisons to shape reward signals • Red-teaming to improve content safety benchmarks