AI Training Specialist (Remote, Contract)
Conducted RLHF preference ranking and prompt evaluation to improve large language model output quality, accuracy, safety, and instruction-following alignment. Performed structured data labeling and annotation to support high-volume evaluation tasks and maintain consistent labeling quality standards. Worked on dataset-facing evaluation workflows to ensure model behavior improvements from labeled feedback. • RLHF preference ranking • Prompt evaluation and scoring • Structured data annotation for training datasets • Quality consistency across high-volume evaluation