AI Trainer & Data Annotator (Freelance, Remote) — RLHF support and alignment evaluation
Supports reinforcement learning from human feedback by contributing human feedback within RLHF training and evaluation workflows. Provides structured human feedback to help models improve alignment with safety, relevance, and quality expectations. Executes quality assurance reviews on feedback-informed outputs to help maintain training integrity. • RLHF workflow participation and human feedback provision • Data validation and feedback quality checks • Alignment-focused content review for safety and relevance • Iterative improvement through evaluation cycles