AI Data Annotator
I served as an AI Training Specialist / Data Annotator executing Reinforcement Learning from Human Feedback (RLHF) for a prominent large language model (LLM) optimization project. The project focused on refining the model's capabilities in handling complex business logic, data structures, and enterprise management applications (such as advanced Excel workflows, SQL query interpretation, and Agile/Jira project frameworks). My day-to-day tasks involved evaluating pairs of AI-generated complex reasoning trajectories, grading them based on strict multi-dimensional rubrics (accuracy, logical flow, formatting constraints, and safety guidelines), and ranking them to train the reward model. A core part of my role was executing fact-checking and grounding procedures to verify technical data and eliminate hallucinations. I regularly drafted detailed analytical rationales (1-2 paragraphs) detailing precisely why a particular model output demonstrated superior logical sequence or constraint adherence, directly contributing to the model's alignment and advanced reasoning capabilities.