Scale AI — AI Data Annotator / AI Trainer
Contributed to artificial intelligence training projects focused on improving the accuracy, reasoning ability, and safety of large language models. The work involved evaluating AI generated responses across medical and general knowledge domains to ensure factual correctness, logical consistency, and adherence to user instructions. Key responsibilities included reviewing and labeling AI outputs for hallucinations, factual errors, reasoning quality, and clinical accuracy. Performed response ranking tasks to identify the most accurate and useful outputs among multiple AI generated answers. Assessed instruction following, completeness, and clarity of responses to support reinforcement learning from human feedback (RLHF) training processes. Additionally contributed to multi-turn conversation evaluation by analyzing dialogue consistency, contextual understanding, and logical flow across interactions. Provided structured feedback to improve model performance, with particular attention to safety, reliability, and real-world applicability of AI systems. This experience strengthened skills in analytical reasoning, attention to detail, structured evaluation, and applied understanding of AI systems in real-world decision making contexts, particularly within healthcare related scenarios.