Skip to content
OpenTrain AIFor AI Companies

Software Engineer, AI Coding Benchmark Development

OpenTrain AI is hiring a senior software engineer to design production-like coding benchmarks and safety tests for frontier coding agents. Contractor role: remote, 20+ hrs/week (min 4 hrs/day), 4–8 week assignment; Python required and 8+ years' experience.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Expert

Experience

Aug 1, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain AI

OpenTrain AI is the centralized platform for people building careers in AI training and data labeling. We help skilled contributors find projects, consolidate their work history, and grow a durable freelance career teaching AI. OpenTrain AI is the hiring organization for this role.

Why AI training work matters

AI training is the human side of creating intelligent systems: people design examples, evaluate outputs, and set safety standards so models behave as intended. This role sits at the intersection of software engineering and AI evaluation — you'll shape how coding agents are measured for both utility and safety.

These projects are typically 100% remote, flexible in schedule, and accessible to experienced engineers who want hands-on, high-impact work improving the reliability and trustworthiness of modern coding assistants.

The role

OpenTrain AI is recruiting a contractor Software Engineer to build and evaluate coding-agent benchmarks used to measure both functional utility and safety/alignment. You will design realistic, production-like tasks, create test suites and reference solutions, and analyze agent rollouts to improve benchmark quality.

  • Contractor role focused on AI coding benchmark development and safety evaluation.
  • Work closely with engineering, QA, and security stakeholders to raise evaluation fidelity.

What you'll do

  • Design realistic coding-agent benchmark tasks using runnable repositories, tests, configs, documentation, and runtime scenarios.
  • Create benign engineering tasks: bug fixes, feature additions, CI repairs, config migrations, integration updates, and runtime-state fixes.
  • Define clear utility requirements that verify correct completion of requested changes.
  • Define safety/alignment requirements that preserve system constraints, developer intent, data integrity, privacy, permissions, and oversight.
  • Develop visible and hidden test suites to evaluate task completion and unsafe agent behavior.
  • Author safe reference solutions and unsafe reference solutions where utility passes but multiple safety/alignment checks fail.
  • Identify and document unsafe shortcuts (e.g., disabling tests, weakening assertions, deleting protected data, leaking secrets, broadening permissions, or bypassing validation).
  • Package tasks with prompts, metadata, runnable repos, evaluators, reference patches, scoring rubrics, and calibration notes.
  • Analyze rollout results from frontier coding agents to assess utility completion and safety/alignment violations.

Requirements

You must meet the core qualifications below; we will verify experience during screening and technical evaluation.

  • Bachelor's or Master's degree in Computer Science or a related technical discipline.
  • Minimum 8 years of hands-on software engineering experience in product companies or startups.
  • Strong programming experience in Python (mandatory).
  • Experience with one or more languages such as Java, Go, or C++.
  • Demonstrated experience building and maintaining production software used by real customers.
  • Practical knowledge of common software vulnerabilities and experience implementing secure code fixes.
  • Experience diagnosing production failures using logs, metrics, and distributed tracing.
  • Strong understanding of software architecture, debugging, performance optimization, and production deployment pipelines.
  • Knowledge of high availability, fault tolerance, scalability, monitoring, logging, and disaster recovery principles.

Helpful background

  • Experience benchmarking or evaluating AI models, especially coding agents.
  • Familiarity with safe vs. unsafe agent behavior analysis and designing hidden adversarial tests.
  • Prior exposure to security assessments, code reviews for vulnerabilities, or secure remediation work.

Commitments & schedule

This is a contractor, part-time assignment with a focused duration. All scheduling and payment are contractor arrangements.

  • Minimum 4 hours per day and at least 20 hours per week.
  • Must provide at least 4 hours of overlap with Pacific Time (PST).
  • Contract duration: approximately 4–8 weeks.
  • Contractor assignment: no paid medical or leave benefits are provided.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Code Evaluation and Benchmarking Engineer

Evaluate and benchmark AI-generated code: review correctness, debug and verify solutions, and build evaluation datasets for frontier models. US-remote, contractor role — 20+ hrs/week (min 4 hrs/day), 1-month contract with 4-hour PST overlap required.

Coding & Software
Text
Remote · United States
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026

AI Benchmark Engineer, Software Engineering

Design and validate multi-agent coding benchmarks using real open-source code changes, Docker, and Python verification scripts. Remote 4-week contractor role for developers in select countries, requiring daily overlap with PST.

Coding & Software
Computer Code Programming
Remote · Bangladesh, Brazil, Colombia +8 more
English
Part-time · Flexible
Entry level

Posted Jul 24, 2026

Software Engineering Code Evaluator

Evaluate and improve AI-generated code while creating benchmarking datasets for advanced software engineering models. This expert, remote contract role is part time at less than 20 hours per week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Expert level

Posted Jul 16, 2026