Ai trainer and evaluator
Project 1: LLM Response Evaluation & Ranking — Outlier AI Worked on a large-scale RLHF pipeline evaluating and ranking responses generated by frontier language models. Tasks involved rating outputs for accuracy, logical coherence, helpfulness, and instruction-following across 1,000+ evaluation tasks. Specialized in STEM-domain prompts covering physics, mathematics, and scientific reasoning — flagging hallucinations, factual errors, and subtle reasoning flaws using structured rubrics.