Skip to content
OpenTrain AIFor AI Companies

SWE-Bench Task Auditor

Audit repository-level software-engineering benchmark tasks that train and evaluate AI models. Review patches, tests, Docker execution, and grading integrity in a remote contract role paying $70 to $90 per hour.

OpenTrain AI

Coding & Software

Remote Hourly · $70–$90/hr

$70–$90/hr

Compensation

1 country

Eligibility

Entry

Experience

Sep 1, 2026

Posted

Open to applicants in

United States

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover specialized projects, build a profile that reflects their experience, and apply in minutes.

As an OpenTrain contractor, you can develop a durable portfolio of AI training work while finding opportunities that match your technical background. Creating an OpenTrain account is free.

About AI Training Work

AI training is the human side of building artificial intelligence. Technical contributors review code, assess model outputs, and evaluate benchmark tasks so AI systems can become more accurate, reliable, and useful.

This work puts experienced software professionals close to the development of cutting-edge AI systems. Remote projects can offer flexible schedules and the opportunity to apply practical engineering judgment to advanced model evaluation.

The Role

OpenTrain is recruiting a SWE-Bench Task Auditor to evaluate repository-level software-engineering benchmark tasks used to train and assess AI models. You will examine task specifications, reference patches, test harnesses, Docker isolation, and grading integrity.

The role combines practical software-engineering judgment with careful evaluation of whether benchmark tasks are correct, reproducible, and resistant to answer leakage or reward hacking. You will provide concise feedback grounded in defined evaluation criteria.

  • Remote contract role for candidates located in the United States
  • Approximately 40 hours per week, with a stated commitment of 20 or more hours weekly
  • Compensation of $70 to $90 per hour
  • Part-time contractor engagement

What You'll Do

You will review software-engineering tasks at the repository level and assess whether they accurately represent the intended work. Your findings will help identify benchmark weaknesses that could distort AI model evaluation.

  • Review repository-level tasks for quality, correctness, and reproducibility
  • Audit reference patches and determine whether they properly address the intended task
  • Inspect test runners, grading behavior, and containerized execution for reliability
  • Assess Docker isolation and overall grading integrity
  • Identify answer leakage, reward hacking, and other evaluation weaknesses
  • Write concise, rubric-based feedback describing findings and recommended improvements

Required Qualifications

This role requires at least three years of professional software-engineering experience, along with meaningful open-source contribution or maintainer experience. The listing is marked entry level, but candidates must meet the stated professional experience and technical requirements.

  • At least three years of professional software-engineering experience
  • Open-source contribution or maintainer experience, including merged pull requests or committer responsibilities
  • Ability to audit reference patches, test runners, Docker isolation, and grading integrity
  • Fluency in Python and at least one of Java, Go, TypeScript, or C++
  • Judgment in detecting answer leakage and reward hacking in software-engineering evaluations
  • Familiarity with SWE-Bench Verified or similar repository-level benchmarks is helpful
  • Prior code-review or task-grading experience is helpful
  • Maintainer history on major Python open-source projects such as Django, Flask, scikit-learn, sympy, or pytest is valuable

Who Should Apply

This opportunity is suited to software engineers who can move comfortably between source code, test infrastructure, containerized execution, and evaluation criteria. It may be especially relevant to open-source contributors and maintainers who understand how repository-level changes should be tested and reviewed.

Strong candidates will be able to explain technical findings clearly, distinguish legitimate task difficulty from benchmark defects, and recognize when evaluation setups create opportunities for leakage or reward hacking.

  • Professional software engineers with strong repository-level debugging judgment
  • Open-source maintainers and contributors with merged pull requests or committer responsibilities
  • Developers experienced with Python and another listed programming language
  • Engineers familiar with benchmark evaluation, code review, or task grading

How to Apply

Create a free OpenTrain account, build your profile around your software-engineering and open-source experience, and apply to this project in minutes. Your OpenTrain profile can help showcase credible AI training and technical evaluation experience as you continue developing your career.

This role requires English-language communication and is available to candidates in the United States. The expected workload is approximately 40 hours per week, with a stated minimum commitment of 20 or more hours weekly.

  • Work location: United States
  • Language: English
  • Work type: Contract and part time
  • Pay: $70 to $90 per hour
  • Expected commitment: 20 or more hours per week

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

AWS Serverless Infrastructure-as-Code Task Auditor

Review AWS serverless architectures and infrastructure-as-code implementations used in AI training and evaluation. Work remotely from the United States for $70-$90 per hour with a 20+ hour weekly commitment.

Coding & Software
Computer Code Programming
Remote · United States
English
Part-time · Flexible
Entry level
Hourly · $70–$90/hr

Posted Sep 1, 2026

Applied Machine Learning Task Auditor

Audit applied machine-learning tasks for sound experiment design, reliable evaluation, and evidence-backed conclusions. This remote US contractor role pays $70-$90 per hour and requires 20+ hours weekly.

Coding & Software
Text
Remote · United States
English
Part-time · Flexible
Entry level
Hourly · $70–$90/hr

Posted Sep 1, 2026

Senior Coding-Agent Benchmark Engineer

Create and evaluate realistic software-engineering benchmarks for coding agents using production-like repositories, secure coding, and rigorous testing. This fully remote contractor assignment runs 4 to 8 weeks.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Sep 2, 2026