AI Model Training & Data Annotation Specialist (Outlier / Scale AI) — Data Annotator / Domain Expert (LLM RLHF)
Contributed to large language model training as a domain expert and data annotator using structured human feedback aligned to task specifications. Evaluated, rated, and ranked model responses for accuracy, coherence, instruction-following, and safety to support RLHF-style fine-tuning and alignment. Performed rubric-based quality scoring and factuality verification to improve downstream model behavior and training data reliability.• Evaluated and graded prompt-and-response pairs using detailed rubrics and safety criteria.• Produced gold-standard/reference answers and refined rationales for consistency and correctness.• Identified hallucinations, edge cases, and model errors with correction notes and documented justifications.• Reviewed peer contributions to ensure adherence to annotation guidelines and SLAs.