Multimodal AI Evaluation Consultant - Independent Contractor
Evaluated multimodal LLM outputs that included electrical diagrams and financial charts for accuracy and reasoning quality. Identified technical inaccuracies and reasoning gaps by cross-checking diagram and financial context. Developed structured evaluation frameworks and prompt-testing methodologies to improve model performance. • Review and assess model outputs involving technical diagrams and financial data • Flag discrepancies, inaccuracies, and reasoning gaps during evaluations • Create evaluation rubrics and structured scoring frameworks • Test and refine prompts using repeatable prompt-testing methods