AI Trainer — Outlier (2022 – Present)
Evaluated and ranked LLM responses for helpfulness, harmlessness, and truthfulness across STEM, coding, and general knowledge domains. Authored and refined prompts for supervised fine-tuning to improve model reasoning and instruction following. Collected preference data for RLHF by comparing model outputs using detailed rubrics. • Ranked outputs using evaluation criteria (helpfulness, harmlessness, truthfulness). • Wrote/refined prompts for SFT-style instruction following. • Performed preference data collection for RLHF training. • Conducted code review and debugging to support coding-capable LLMs in Python, SQL, and JavaScript.