Skip to content
OpenTrain AIFor AI Companies

Coding-Agent Benchmark Engineer

Design and evaluate realistic coding-agent tasks using production-like repositories, tests, evaluators, and scoring rubrics. This contract role is open worldwide, requires English fluency and 20+ hours per week, and calls for deep software engineering experience.

Apply now
OpenTrain AI

Coding & Software

100% Remote

Contract, part-time

Engagement

Remote

Location

Oct 7, 2026

Posted

Open worldwide

The work

You will create and evaluate software engineering tasks that measure whether coding agents can complete work while preserving safety, privacy, permissions, data integrity, and developer intent. The work uses production-like repositories, runtime scenarios, configurations, documentation, tests, evaluators, reference patches, and scoring rubrics.

  • Build benchmark tasks for bug fixes, feature additions, CI repairs, configuration migrations, integration updates, and runtime-state fixes.
  • Define utility and safety requirements, then create visible and hidden tests for each task.
  • Develop safe and intentionally unsafe reference solutions for evaluation scenarios.
  • Assess shortcuts such as disabling tests, weakening assertions, deleting protected data, leaking secrets, broadening permissions, or bypassing validation.
  • Package prompts, metadata, runnable repositories, evaluators, reference patches, rubrics, and calibration notes.
  • Analyze coding-agent rollouts for utility and safety violations.
  • Work with engineering, quality assurance, and security stakeholders to improve evaluator reliability and benchmark difficulty.

What it pays and takes

Pay details are not provided in the listing. The role is structured as part-time contract work and requires at least 20 hours each week.

  • Open worldwide.
  • English fluency required.
  • At least eight years of hands-on software engineering experience in leading product companies or technology startups.
  • Bachelor's or master's degree in computer science or a related technical field.
  • Mandatory Python expertise, plus proficiency in Java, Go, or C++.
  • Experience building and maintaining production software used by real customers.
  • Strong skills in software architecture, debugging, performance optimization, deployment pipelines, monitoring, logging, and distributed tracing.
  • Knowledge of high availability, fault tolerance, scalability, disaster recovery, secure code review, vulnerability remediation, and complex production failure analysis.

How it works

Apply on OpenTrain with your resume, then complete the application on the hiring site.

About AI training work

AI training work is the human work behind modern artificial intelligence, including preparing examples, evaluating model behavior, and testing whether systems follow instructions. Experienced software engineers are needed for specialist projects such as coding-agent evaluation, where technical judgment helps measure quality, reliability, and safe behavior.

Requirements

  • Experience: Entry level
  • Languages: English

How to apply

  1. Apply here on OpenTrain. You create a free account, and we send you to the hiring platform.
  2. Complete your application on the hiring platform.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar open roles

View all AI training jobs

Bioinformatics AI Evaluation Task Designer

Design rigorous bioinformatics and computational genomics tasks that test whether AI models can analyze data, write code, and produce verifiable scientific results. This remote five-week contractor assignment pays $150 per approved task.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Per task · $150/label

Posted Oct 7, 2026

GenAI Security Evaluation Engineer

Build and test vulnerable GenAI agent and RAG codebases, annotate security risks, and evaluate detection tools at $150 per hour. This contract role requires 20+ hours per week and strong application security experience.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $150/hr

Posted Oct 7, 2026

Backend AI Coding Task Creator

Create realistic backend coding tasks and reliable verifiers for AI systems. This contractor role offers a flexible schedule of about 15 hours per week and pays $30 to $100 per hour equivalent, paid per task that meets project specifications.

Coding & Software
Computer Code Programming
Remote · United Arab Emirates, Argentina, Austria +50 more
English
Part-time · Flexible
Mid-Senior level
Per task · $30–$100/hr equivalent

Posted Oct 6, 2026

Python Backend Developer for AI Agent Data

Build and test Python backend connectors or create realistic AI agent tasks and evaluation rubrics. This remote, four-week contract pays $300 per approved task and requires at least 20 hours per week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Per task · $300/label

Posted Oct 4, 2026

Backend Coding Task Creator and AI Evaluator

Create realistic backend coding tasks, build verifiers, and evaluate AI-generated solutions for correctness, performance, testing, and maintainability. Pay is $30 to $100 per hour equivalent, paid per task that meets the project specifications.

Coding & Software
Computer Code Programming
Remote · United Arab Emirates, Argentina, Austria +50 more
English
Part-time · Flexible
Mid-Senior level
Per task · $30–$100/hr equivalent

Posted Oct 3, 2026