AI Visual Reasoning Evaluator & Reviewer (Project Escher) — Handshake AI (May 2026 – Present)
Designed and reviewed AI visual-reasoning evaluation tasks to measure frontier vision-language model performance. Created test cases featuring charts, diagrams, maps, and infographics targeting spatial reasoning, color discrimination, counting, and partial occlusion. Reviewed peer submissions using a 5-point rubric and flagged low-quality answers early to reduce rework. • Targeted weaknesses such as hidden/overlapping shapes, fine color differences, and multi-panel tracking • Promoted to reviewer after sustained quality • Graded submissions on answer correctness, prompt clarity, copyright, and model failure modes • Achieved elevated model-failure rates by concentrating on known weak spots