AI/ML Trainer at Outlier AI (Mar. 2024 – Present)
Trained multiple large language models and generative AI systems using reinforcement learning from human feedback (RLHF) and semi-supervised learning (SSL). Focused on preparing and applying human feedback signals to improve model behavior and downstream responses. Performed iterative training workflows typical of LLM fine-tuning and evaluation cycles. • RLHF and SSL-based training activities • Working with LLM and GenAI training datasets/workflows • Iterative improvement using human feedback • Model training and performance refinement