Evaluate real software repositories and help measure how well AI coding systems fix bugs. This remote, three-month contractor assignment offers 20 hours per week for engineers skilled in codebases, testing, Git, and Docker.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover projects, build a professional profile, and apply for opportunities that match their skills.
OpenTrain AI is hiring and contracting for this remote software engineering evaluation assignment.
About AI Training and Software Evaluation
AI training is the human side of building artificial intelligence. For coding systems, experienced software professionals help evaluate generated code, test whether bugs are fixed correctly, and create reliable benchmarks from real-world projects.
This work contributes to the development of advanced AI systems while offering flexible remote opportunities for people with software engineering experience.
The Role
As an LLM Evaluation & Repository Validation Engineer, you will help build verifiable software engineering tasks from public repository histories. You will analyze GitHub issues, configure code repositories, establish development environments, review test coverage and code quality, and run projects locally to assess bug-fixing performance.
You will also help identify repositories and issues that are especially challenging for large language models, contributing technical judgment from a software engineering perspective.
Contractor assignment
20 hours per week
Three-month engagement
Remote worldwide
English-language work
What You'll Do
You will work with open-source software repositories and collaborate with researchers on repository selection and evaluation task design. The work requires hands-on technical investigation rather than purely theoretical review.
Analyze and triage GitHub issues across open-source libraries.
Set up and configure repositories, including Dockerization and development environments.
Evaluate unit test coverage and code quality.
Run, modify, and test real-world projects locally.
Assess LLM performance on software engineering and bug-fixing tasks.
Identify repositories and issues that present meaningful challenges for LLMs.
Collaborate with researchers on repository selection and task design.
Apply technical judgment from a senior software engineering perspective.
Requirements
This role is suited to someone who can confidently navigate complex codebases, debug software locally, and assess whether tests and fixes are reliable. Strong experience with public GitHub repositories is important.
Strong experience with at least one of Python, JavaScript, Java, Go, Rust, C, C++, C#, or Ruby.
Hands-on proficiency with Git, Docker, and basic software pipeline setup.
Ability to navigate complex codebases and debug locally.
Ability to run, modify, and test projects locally.
Experience evaluating unit tests, debugging code, and assessing bug-fixing performance.
Strong experience working with public GitHub repositories and complex codebases.
Helpful Background
The following experience is helpful for contributing effectively to repository validation and AI evaluation work, but is not presented as a mandatory requirement.
Tech lead-level software engineering experience.
Experience contributing to or evaluating open-source projects.
Familiarity with high-quality public GitHub repositories, including widely used repositories with 500 or more stars.
Previous work in LLM research or evaluation.
Experience with developer tools and automation agents.
Comfort collaborating with researchers on AI evaluation tasks.
Remote collaboration experience with overlapping Pacific Time hours.
Build Your AI Training Career with OpenTrain
AI training and data-labeling work spans code evaluation, model feedback, language, images, audio, and other forms of human-reviewed data. Contributors use their professional knowledge to help shape how modern AI systems perform.
Through OpenTrain, you can build a profile that showcases relevant experience, discover projects aligned with your skills, and develop a lasting portfolio in this rapidly growing field.
Work remotely with a flexible part-time schedule.
Apply software engineering expertise to cutting-edge AI evaluation.
Build credible experience in coding and LLM assessment.
Create a stronger profile for future AI training opportunities.
Evaluate real Rust repositories and GitHub issues to help measure how LLMs solve software bugs. This part-time contractor role offers 20+ hours per week for engineers experienced with Rust, Git, Docker, and local testing.
Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.
Join OpenTrain to build LLM evaluation and training datasets by validating real open-source codebases with C#. This remote contractor role requires 3+ years of software engineering, 20+ hours/week (options to 30–40) and is open to candidates in specified countries.