Audit repository-level software-engineering benchmark tasks that train and evaluate AI models. Review patches, tests, Docker execution, and grading integrity in a remote contract role paying $70 to $90 per hour.
Coding & Software
Remote Hourly · $70–$90/hr
$70–$90/hr
Compensation
1 country
Eligibility
Entry
Experience
Sep 1, 2026
Posted
Open to applicants in
United States
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover specialized projects, build a profile that reflects their experience, and apply in minutes.
As an OpenTrain contractor, you can develop a durable portfolio of AI training work while finding opportunities that match your technical background. Creating an OpenTrain account is free.
About AI Training Work
AI training is the human side of building artificial intelligence. Technical contributors review code, assess model outputs, and evaluate benchmark tasks so AI systems can become more accurate, reliable, and useful.
This work puts experienced software professionals close to the development of cutting-edge AI systems. Remote projects can offer flexible schedules and the opportunity to apply practical engineering judgment to advanced model evaluation.
The Role
OpenTrain is recruiting a SWE-Bench Task Auditor to evaluate repository-level software-engineering benchmark tasks used to train and assess AI models. You will examine task specifications, reference patches, test harnesses, Docker isolation, and grading integrity.
The role combines practical software-engineering judgment with careful evaluation of whether benchmark tasks are correct, reproducible, and resistant to answer leakage or reward hacking. You will provide concise feedback grounded in defined evaluation criteria.
Remote contract role for candidates located in the United States
Approximately 40 hours per week, with a stated commitment of 20 or more hours weekly
Compensation of $70 to $90 per hour
Part-time contractor engagement
What You'll Do
You will review software-engineering tasks at the repository level and assess whether they accurately represent the intended work. Your findings will help identify benchmark weaknesses that could distort AI model evaluation.
Review repository-level tasks for quality, correctness, and reproducibility
Audit reference patches and determine whether they properly address the intended task
Inspect test runners, grading behavior, and containerized execution for reliability
Assess Docker isolation and overall grading integrity
Identify answer leakage, reward hacking, and other evaluation weaknesses
Write concise, rubric-based feedback describing findings and recommended improvements
Required Qualifications
This role requires at least three years of professional software-engineering experience, along with meaningful open-source contribution or maintainer experience. The listing is marked entry level, but candidates must meet the stated professional experience and technical requirements.
At least three years of professional software-engineering experience
Open-source contribution or maintainer experience, including merged pull requests or committer responsibilities
Ability to audit reference patches, test runners, Docker isolation, and grading integrity
Fluency in Python and at least one of Java, Go, TypeScript, or C++
Judgment in detecting answer leakage and reward hacking in software-engineering evaluations
Familiarity with SWE-Bench Verified or similar repository-level benchmarks is helpful
Prior code-review or task-grading experience is helpful
Maintainer history on major Python open-source projects such as Django, Flask, scikit-learn, sympy, or pytest is valuable
Who Should Apply
This opportunity is suited to software engineers who can move comfortably between source code, test infrastructure, containerized execution, and evaluation criteria. It may be especially relevant to open-source contributors and maintainers who understand how repository-level changes should be tested and reviewed.
Strong candidates will be able to explain technical findings clearly, distinguish legitimate task difficulty from benchmark defects, and recognize when evaluation setups create opportunities for leakage or reward hacking.
Professional software engineers with strong repository-level debugging judgment
Open-source maintainers and contributors with merged pull requests or committer responsibilities
Developers experienced with Python and another listed programming language
Engineers familiar with benchmark evaluation, code review, or task grading
How to Apply
Create a free OpenTrain account, build your profile around your software-engineering and open-source experience, and apply to this project in minutes. Your OpenTrain profile can help showcase credible AI training and technical evaluation experience as you continue developing your career.
This role requires English-language communication and is available to candidates in the United States. The expected workload is approximately 40 hours per week, with a stated minimum commitment of 20 or more hours weekly.
Review AWS serverless architectures and infrastructure-as-code implementations used in AI training and evaluation. Work remotely from the United States for $70-$90 per hour with a 20+ hour weekly commitment.
Audit applied machine-learning tasks for sound experiment design, reliable evaluation, and evidence-backed conclusions. This remote US contractor role pays $70-$90 per hour and requires 20+ hours weekly.
Create and evaluate realistic software-engineering benchmarks for coding agents using production-like repositories, secure coding, and rigorous testing. This fully remote contractor assignment runs 4 to 8 weeks.