AI Data Trainer – Independent Contractor
Evaluated and refined generative AI model outputs using reinforcement learning from human feedback, focusing on improving response quality and alignment. Conducted structured prompt-response validation, hallucination detection, and generated feedback reports to support iterative model development. Maintained high annotation accuracy and collaborated on systematic QA initiatives targeting reasoning quality and diverse output categories. • Applied RLHF techniques to large-scale generative AI tasks • Performed multi-step quality assurance and reasoning performance analysis • Produced structured, team-adopted feedback documentation • Drove improvement on accuracy, coherence, and safety benchmarks.