AI Trainer (Remote/Contract)
Created high-quality training datasets for large language models (LLMs), including prompt-response pairs, reasoning chains, and domain-specific content across IT, software development, and general knowledge. Evaluated and ranked AI-generated outputs using RLHF (Reinforcement Learning from Human Feedback) to improve model accuracy and alignment. Identified model failure modes, edge cases, and hallucinations, providing structured feedback to guide model fine-tuning efforts.• Built training examples comprising prompt-response and reasoning content.• Performed evaluation and ranking of model outputs for reward learning.• Produced structured error analysis for hallucinations and edge cases.• Assisted in developing annotation guidelines to keep labels consistent across contributors.