AI Trainer & Data Annotator at Outlier AI
Performed RLHF-based (Reinforcement Learning from Human Feedback) data annotation and response assessment to support model training and optimisation workflows at scale. Evaluated and ranked AI-generated responses for quality to directly contribute to large language model improvement. Reviewed and validated annotated datasets using structured quality assurance processes while maintaining annotation consistency under evolving evaluation criteria. • RLHF data annotation and response assessment. • Quality evaluation and ranking of AI-generated responses. • Dataset validation with QA checks. • Policy-guideline compliance in fast-paced tasks.