AI Evaluator - Handshake
Analyzed and evaluated domain-specific prompts to assess large language model performance for image recognition across multiple specialized subfields. Reviewed model outputs for scientific accuracy, clarity, and depth, and delivered expert feedback to improve AI understanding of complex visualizations. Conducted independent research to support prompt development and evaluation workflows for ongoing quality improvement. • Developed prompt evaluation workflows for LLM image recognition tasks • Assessed outputs against criteria for accuracy and clarity • Provided expert review to refine model behavior and responses • Supported iterative research efforts to improve evaluation quality