AI & Language Specialist (Remote/Freelance) — AI Training & Annotation for LLMs
Contributed to Large Language Model development using Reinforcement Learning from Human Feedback (RLHF), focusing on improving response accuracy and safety. Created, evaluated, and refined question–answer pairs to simulate real-world scenarios for supervised alignment. Performed iterative review to support model behavior that matches user intent and reduces issues such as hallucinations or bias. • Question–answer pair creation and refinement • RLHF-focused response safety and accuracy checks • Evaluation of AI outputs for quality and coherence • Alignment with real-world user intent in simulated scenarios