Freelance AI Trainer & Dataset Engineer (Remote) — RLHF training datasets, preference labeling, and LLM evaluation
Led RLHF-oriented dataset engineering by preparing preference-labeled and QA-reviewed samples for supervised fine-tuning and LLM evaluation. Built a repeatable annotation pipeline spanning transcription-derived text, preference labeling, and quality assurance to minimize pipeline errors. Evaluated LLM outputs for factual accuracy, response quality, and safety using preference ranking criteria aligned with RLHF methodologies. • Prepared and annotated 1,000+ labeled text samples for supervised fine-tuning. • Applied multi-pass QA reviews to deliver zero-defect datasets. • Implemented preference labeling workflows and quality checks across model cycles. • Performed LLM output evaluation using RLHF-aligned ranking criteria.