Freelance AI Data Annotator & AI Trainer (RLHF)
Executed RLHF workflows by ranking AI-generated responses and assigning scalar preference feedback to guide LLM fine-tuning. Wrote preference annotations and structured justifications to support training signal quality. Performed evaluation-style annotation to ensure outputs met client instruction-following, accuracy, and safety expectations. • Response ranking and preference annotation for RLHF • Scalar feedback and detailed annotation rationale • Consistency checks against client requirements and style guides • Guideline-contribution via iterative reporting