AI Specialist — Handshake AI (Remote)
Evaluated AI-generated outputs across multiple task formats to improve model accuracy, consistency, and reliability. Conducted structured evaluation (eval) tasks to assess AI system performance against established quality standards. Identified response errors, inconsistencies, and edge cases, providing detailed feedback to support model improvement. • Reviewed complex AI-generated content • Validated output quality using analytical and critical-thinking methods • Performed data-driven assessments and performance evaluations for ongoing training and optimization • Fed back findings to improve model behavior and reduce failure modes