Senior AI Annotation & Data Labeling Specialist at Scale AI (Remote)
Performed RLHF workflows to compare model responses, rank preferences, and assign quality scores while checking policy compliance. Annotated and evaluated LLM outputs for reasoning, factuality, summarization, instruction-following, and conversational behavior. Audited datasets for consistency, formatting compliance, and guideline adherence across high-volume production projects. • Response comparison and preference ranking • Quality scoring, policy evaluation, and safety checks • Hallucination and unsafe-response detection • Calibration exercises and contributor QA feedback