Data Labelling Specialist — Open Train AI (Remote)
Annotated and labeled text, image, and audio samples to support LLM fine-tuning pipelines with response ranking and scoring. Applied RLHF techniques to align AI-generated responses using human feedback-based evaluation signals. Maintained high annotation consistency across multi-annotator workflows and improved alignment quality through metric-driven review. • Labelled 50,000+ samples for LLM fine-tuning. • Ranked and scored AI responses for alignment training. • Sustained 96%+ consistency scores while reducing inter-rater disagreement by 18%. • Developed and refined labeling guidelines and ontologies with ML engineers.