AI Training Contributor | Handshake
Evaluated large language model responses across business writing, marketing, sales communication, general knowledge, and life sciences using rubric-based scoring for factual accuracy, instruction-following, helpfulness, and safety. Performed side-by-side output comparisons to support reinforcement learning from human feedback (RLHF) workflows and selected preferred responses. Identified and flagged hallucinations, factual errors, unsupported claims, and subtle instruction-adherence failures. • Rubric-based factuality grading and helpfulness evaluation • Side-by-side model comparison for RLHF preference judgments • Hallucination detection and safety/instruction adherence checks • Written rationales documenting ranking decisions and edge cases