AI Training Data Contributor & RLHF Evaluator | DataAnnotation.tech
Evaluated and rewrote 900+ prompt-response pairs for LLM RLHF training. Assessed outputs for factual accuracy, tone, helpfulness, and instruction adherence, then rewrote 400+ low-quality responses to improve clarity and contextual correctness. Maintained a 91% evaluator agreement score with expert reviewers and identified systematic model failure patterns to refine prompt guidelines. • RLHF output evaluation (accuracy/tone/helpfulness/instruction) • Prompt-response rewriting and style/tone improvements • Inter-evaluator agreement calibration • Failure-mode analysis and edge-case escalation