Skip to content
OpenTrain AIFor AI Companies

Secure Coding-Agent Benchmark Engineer

Design realistic coding-agent benchmark tasks using production-like software, security tests, and evaluation tools. This remote contractor role requires strong Python skills, deep software engineering experience, and at least 20 hours per week.

Apply now
OpenTrain AI

Coding & Software

100% Remote

Contract, part-time

Engagement

Remote

Location

Oct 9, 2026

Posted

Open worldwide

The work

You will design and build realistic benchmark tasks that measure whether AI coding agents can complete software work safely and correctly. Tasks use production-like repositories, tests, configurations, documentation, and runtime scenarios.

You will combine software engineering judgment with structured evaluation of coding-agent behavior. The work includes building tasks, testing agent solutions, analyzing failures, and improving benchmark reliability and difficulty.

  • Create benchmark tasks for bug fixes, feature additions, CI repairs, configuration migrations, integration updates, and runtime-state fixes.
  • Define utility and safety requirements that explain what an agent must do and which unsafe shortcuts it must avoid.
  • Develop visible and hidden test suites, safe and unsafe reference solutions, scoring rubrics, calibration notes, and runnable evaluators.
  • Review agent rollouts for problems such as disabling tests, weakening assertions, deleting protected data, leaking secrets, broadening permissions, or bypassing validation.
  • Package benchmark tasks and work with engineering, quality assurance, and security stakeholders to improve evaluator reliability and difficulty.

What it pays and takes

This is a remote contractor assignment expected to last about 4 to 8 weeks. The work is listed as part-time, with a commitment of at least 20 hours per week and regular overlap with Pacific Time.

The role requires advanced hands-on software engineering experience and secure development knowledge. Helpful experience includes creating coding-agent benchmarks, evaluating coding-agent rollouts, or designing safety and alignment tests.

  • Pay: Not provided in the listing.
  • Schedule: At least 20 hours per week, including at least 4 hours per day and 4 hours of overlap with Pacific Time.
  • Location: Fully remote and open worldwide.
  • Language: Fluent in English.
  • Experience: At least 8 years of hands-on software engineering experience in leading product companies or technology startups.
  • Education: A bachelor's or master's degree in computer science or a related technical discipline.
  • Programming: Strong Python expertise is mandatory, plus proficiency in Java, Go, or C++.
  • Production systems: Experience building and operating software used by real customers at scale, including deployment pipelines, monitoring, logging, observability, distributed tracing, high availability, fault tolerance, scalability, and disaster recovery.
  • Security: Practical experience identifying vulnerabilities, implementing secure fixes, and validating remediation. Familiarity with injection, authorization, memory safety, race conditions, deserialization, secrets management, and input validation issues is required.

How it works

Apply on OpenTrain with your resume and then complete the application on the hiring site.

About AI training work

AI training work uses human-created examples, reviews, and evaluations to improve how artificial intelligence systems behave. This role focuses on testing coding agents, and experienced software engineers are needed to judge whether their solutions are useful, secure, and aligned with developer intent.

Requirements

  • Experience: Entry level
  • Languages: English

How to apply

  1. Apply here on OpenTrain. You create a free account, and we send you to the hiring platform.
  2. Complete your application on the hiring platform.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar open roles

View all AI training jobs

Earth Sciences AI Task Development Expert

Create realistic Earth sciences tasks for AI training and evaluation using scientific code, environmental datasets, simulations, and automated tests. This remote contractor role is open worldwide and requires fluent English and advanced technical experience.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Oct 9, 2026

Electrical Engineering and Computer Science AI Quality Lead

Review code, simulations, circuit designs, automated tests, and technical AI evaluation tasks while guiding a pod of 5 to 10 trainers. This four-week contractor assignment requires advanced technical experience, Linux skills, and at least 20 hours per week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Oct 9, 2026

Chemical Engineering AI Training Pod Lead

Lead a pod of technical trainers creating and reviewing chemical engineering tasks for AI training. Evaluate simulations, models, graders, and engineering results using advanced technical judgment.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Oct 9, 2026

Civil And Structural Engineering Task Quality Lead

Review civil and structural engineering models, simulations, reference solutions, and automated graders for AI evaluation tasks. This remote contractor assignment requires advanced engineering experience, programming skills, and at least 20 hours per week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Oct 9, 2026

Agentic AI Software Engineer

Build MCP tools, improve AI coding assistant workflows, and create production-ready data applications for agentic AI systems. This part-time contract role is open worldwide and requires 20 or more hours each week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Oct 9, 2026