AI Evaluation and Training Specialist (Contract) - Handshake AI
Provided evaluation and testing of multimodal AI systems across text, image, video, and audio modalities to improve model performance and reliability. Developed adversarial prompts and benchmark scenarios to surface hallucinations, reasoning gaps, and edge-case failures in frontier models. Applied structured rubrics within certification-gated workflows to assess complex datasets and generate actionable findings for training and iteration. • Designed multi-hop and visual-reasoning prompts to stress-test model outputs • Evaluated entity tagging, caption accuracy, and transcription quality using detailed scoring rubrics • Wrote model feedback and prompt specifications to support reproducible evaluation workflows • Translated technical evaluation criteria into clear documentation for certification processes