AI Trainer & RLHF Evaluator — Outlier.ai (Remote)
Performed RLHF preference-ranking tasks by evaluating pairs of AI-generated responses for helpfulness, accuracy, safety, and coherence to support reward model training.• Wrote diverse, high-quality prompts to expand LLM training datasets across creative writing, general knowledge, logical reasoning, and real-world problem-solving.• Completed chain-of-thought annotation by verifying mathematical and logical reasoning steps within AI outputs to improve transparency and accuracy.• Conducted safety evaluation by flagging harmful, biased, or misleading content and maintained above-threshold quality scores for premium task access.