Skip to content
OpenTrain AIFor AI Companies

AI Benchmark Engineer — Software Engineering

Design and validate multi-agent coding benchmarks using real open-source code changes, Docker, and Python verification scripts. Remote 4-week contractor role for developers in select countries, requiring daily overlap with PST.

OpenTrain AI

Coding & Software

Remote

11 countries

Eligibility

Entry

Experience

Jul 24, 2026

Posted

Open to applicants in

Bangladesh Brazil Colombia Egypt Ghana India Indonesia Kenya Nigeria Türkiye Vietnam

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people start and grow careers teaching AI by discovering projects, building a unified profile, and applying quickly to roles that match their skills.

OpenTrain connects experienced contributors and newcomers alike to meaningful AI-training work — the human effort that shapes how modern AI systems behave.

About AI training work

AI training (data labeling, annotation, and human feedback) is the human side of building AI. Contributors annotate code, write and evaluate model outputs, and create benchmarks that teach and test models.

This work is remote, flexible, often accessible without prior labeling experience, and places you on the cutting edge of how state-of-the-art AI systems are built and evaluated.

The role — what an AI Benchmark Engineer does

You will design and build multi-agent benchmark tasks derived from real open-source code changes (bug fixes, migrations, refactors) and ensure those tasks run reproducibly inside containerized evaluation environments.

The role centers on precise technical specifications, Python-based verification, Dockerized task execution, and decomposing complex edits across coordinated sub-agents.

What you'll do day-to-day

  • Build multi-agent benchmark tasks grounded in real open-source code changes, including bug fixes, migrations, and refactors.
  • Work with the Harbor evaluation framework to run and validate tasks inside Docker environments.
  • Write clear, precise task instructions specifying file paths, function signatures, expected behavior, and constraints.
  • Design and implement Python verification scripts to validate correctness of agent-generated code changes.
  • Create decomposition strategies that split complex code changes across independent sub-agents.
  • Run, debug, and refine tasks within containers to ensure reproducibility and determinism.
  • Evaluate task performance signals and iterate on task quality, clarity, and difficulty.

Requirements

This role requires solid practical engineering experience and familiarity with evaluation tooling and workflows.

  • 5+ years of experience in Python and JavaScript development.
  • Experience with AI coding benchmarks (for example, SWE-bench, Terminal-Bench).
  • Strong experience reading and navigating large open-source codebases (Django, Flask, FastAPI, Node.js, or similar).
  • Familiarity with Git workflows: pull requests, diffs, cherry-picking, and working with specific commits.
  • Comfortable working with Docker, including writing Dockerfiles, building images, and debugging containers.
  • Experience writing test scripts (pytest, unittest, or custom assertion-based testing).
  • Ability to write clear, precise, and unambiguous technical specifications.

Helpful background

  • Interest in AI agentic behavior and multi-agent coordination.
  • Familiarity with LLM evaluation and reasoning benchmarks is a plus.

Location, schedule, and contract details

This is a remote contractor role open to contributors located in Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Kenya, Nigeria, Turkey, or Vietnam.

Time expectations: 20+ hours per week, with a stated requirement of 8 hours per day and a 4-hour overlap with Pacific Standard Time (PST). Duration: 4-week contract. This contractor position does not include medical or paid leave.

Who should apply and how it works

Apply if you are an experienced Python/JavaScript engineer who enjoys designing reproducible evaluation tasks and working with containers and verification tooling. This role suits engineers who can read large codebases and translate changes into precise, testable tasks.

To apply, create or use your OpenTrain account, complete your profile, and submit your application. OpenTrain manages hiring and contracting for this role.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

SWE Bench Data Engineer / Data Scientist — Benchmark Evaluation

Join OpenTrain AI as a contract, part-time SWE Bench contributor working 20+ hours/week on benchmark-driven data engineering and data-science evaluation. Use Python to build reproducible pipelines, process structured and unstructured data, and validate real-world workflows (English required).

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +7 more
English
Part-time · Flexible
Intermediate level

Posted Jul 16, 2026

Machine Learning Benchmark Evaluator

Experienced ML engineers wanted for part-time, remote contract work evaluating production-grade model training, evaluation, and inference pipelines; requires 3+ years of ML engineering experience, strong Python, and PyTorch/TensorFlow/JAX familiarity. 20+ hours/week; apply through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +7 more
English
Part-time · Flexible
Expert level

Posted Jul 17, 2026

Machine Learning Engineer, Benchmarking & Evaluation

Join OpenTrain as a remote Machine Learning Engineer focused on benchmark-driven evaluation of real-world ML systems. This contractor role requires 3+ years of ML engineering experience, strong Python skills, and availability 20+ hrs/week with PST overlap.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +7 more
English
Part-time · Flexible
Intermediate level

Posted Jul 16, 2026