DataAnnotation, AI Trainer (Remote, USA)
Worked as an AI Trainer using Reinforcement Learning from Human Feedback (RLHF) to help machine learning models improve their behavior and performance. Reviewed others’ work to perform quality assurance and ensure outputs complied with legal and ethical obligations. Provided training feedback aligned with model improvement goals and policy compliance requirements. • Used RLHF-based training workflows • Performed peer review for quality assurance • Ensured legal/ethical compliance during training • Supported iterative model improvement through feedback