AI Expert Contributor (DevOps) at Snorkel AI (benchmarking and evaluation)
Contributed to Snorkel AI as an AI Expert Contributor (DevOps) by building and verifying software engineering benchmarks to stress-test frontier large language models. Focused on deterministic evaluation, reproducibility, and exhaustive coverage of technical edge cases during benchmark automation and validation. Investigated model failure modes and performance metrics to iteratively improve verification rigor and evaluation frameworks. • Engineered benchmark pipelines using Python and system architectures to evaluate LLM behavior • Created automated validation suites emphasizing determinism and reproducibility • Analyzed LLM failure modes and performance metrics to strengthen evaluation coverage • Supported verification and benchmarking infrastructure as part of DevOps work for AI models