Software Engineer — AI Code Ranking
This project evaluates and compares AI-generated code-repository responses for software engineering tasks. Annotators review two model-generated solutions to the same GitHub/repository issue, analyze their reasoning process, code changes, tool usage, correctness, and overall implementation quality, then classify strengths, weaknesses, and overall preference using a structured evaluation rubric. The goal is to measure how effectively AI systems can investigate, modify, and reason about real-world codebases while producing maintainable and correct solutions.