Senior Data Scientist — LLM Evaluation & AI Systems (Abundant)
Led LLM benchmarking and evaluation programs by creating structured test suites spanning reasoning, multimodal understanding, and domain-specific tasks for Anthropic, OpenAI, and Gemini. Built automated evaluation pipelines to compare model outputs at scale and surface hallucination patterns, reasoning inconsistencies, and edge-case failure modes. Designed case-study tasks that stress-test multimodal reasoning and statistical inference across frontier models. • Developed structured evaluation frameworks for reliability and trust • Authored and executed large-scale test suites for provider model comparisons • Analyzed failure modes and converted results into actionable research/product insights • Coordinated with cross-functional teams on scalable evaluation methodology