AI Benchmark Creation & Code Evaluation
Created and evaluated software engineering benchmark tasks for training frontier AI models at AfterQuery, a YC-backed AI research lab. Tasks involved identifying bugs in Python and JavaScript codebases, writing test suites to verify correct solutions, and building broken Docker environments that required multi-step terminal debugging to fix. Evaluated AI-generated code solutions for correctness, edge case handling, and adherence to software engineering best practices.