AI Data Trainer & Annotator | Outlier AI (Remote)
Served as an AI Data Trainer and Annotator, producing labeled datasets and high-quality quality checks for LLM fine-tuning tasks. Provided comparative human-style feedback on AI-generated outputs to support RLHF pipelines and improve response quality. Built and applied annotation guidelines while identifying edge cases and model errors to inform iterative retraining decisions. • Annotated and quality-checked 10,000+ text/code/reasoning samples for LLM fine-tuning • Delivered structured comparative feedback for RLHF to improve response quality by ~18% • Developed annotation guidelines and style guides adopted by 12 annotators • Reported edge cases and systematic errors; collaborated to refine HHH evaluation rubrics