Project Escher
Author unambiguous adversarial image-based prompts designed to identify where vision-language models fail. Source reference images, construct prompts with precise ground-truth expectations, and iteratively refine wording to remove ambiguity while preserving the failure-inducing edge case. Evaluate model responses against the authored ground truth and document failure modes for downstream model training.