Data Annotation Tech
Conducted high-level Reinforcement Learning from Human Feedback (RLHF) to evaluate, train, and optimize advanced Large Language Models (LLMs). My day-to-day responsibilities included drafting complex, multi-turn adversarial prompts ('red-teaming') to test model boundaries, and side-by-side comparative analysis of model outputs. I evaluated responses based on rigorous multi-axis criteria, including factual accuracy, structural logic, tone, helpfulness, and adherence to constraints. A significant portion of my work involved deep-dive factual verification—independently researching and cross-referencing claims to ensure absolute accuracy. When model outputs fell short, I wrote high-quality, ground-truth target responses and provided detailed, text-based justifications explaining exactly why one response outperformed another based on formatting, logical coherence, and prompt alignment.