Principal AI Research Scientist and LLM Evaluation Lead – Open Mind Research Institute (Remote)
Led and managed LLM evaluation work for the OpenMind Computational Correctness Benchmark. Focused on designing evaluation frameworks and coordinating asynchronous research operations across multiple time zones. Oversaw delivery of evaluation methods intended to assess model computational correctness and related behaviors. • Directed a team of 8 senior researchers on evaluation framework design. • Managed remote research workflow across 5 time zones. • Conducted prompt/response assessment activities as part of LLM evaluation. • Supported benchmark engineering and correctness-focused evaluation deliverables.