You will design and build realistic benchmark tasks that measure whether AI coding agents can complete software work safely and correctly. Tasks use production-like repositories, tests, configurations, documentation, and runtime scenarios.
You will combine software engineering judgment with structured evaluation of coding-agent behavior. The work includes building tasks, testing agent solutions, analyzing failures, and improving benchmark reliability and difficulty.
- Create benchmark tasks for bug fixes, feature additions, CI repairs, configuration migrations, integration updates, and runtime-state fixes.
- Define utility and safety requirements that explain what an agent must do and which unsafe shortcuts it must avoid.
- Develop visible and hidden test suites, safe and unsafe reference solutions, scoring rubrics, calibration notes, and runnable evaluators.
- Review agent rollouts for problems such as disabling tests, weakening assertions, deleting protected data, leaking secrets, broadening permissions, or bypassing validation.
- Package benchmark tasks and work with engineering, quality assurance, and security stakeholders to improve evaluator reliability and difficulty.
What it pays and takes
This is a remote contractor assignment expected to last about 4 to 8 weeks. The work is listed as part-time, with a commitment of at least 20 hours per week and regular overlap with Pacific Time.
The role requires advanced hands-on software engineering experience and secure development knowledge. Helpful experience includes creating coding-agent benchmarks, evaluating coding-agent rollouts, or designing safety and alignment tests.
- Pay: Not provided in the listing.
- Schedule: At least 20 hours per week, including at least 4 hours per day and 4 hours of overlap with Pacific Time.
- Location: Fully remote and open worldwide.
- Language: Fluent in English.
- Experience: At least 8 years of hands-on software engineering experience in leading product companies or technology startups.
- Education: A bachelor's or master's degree in computer science or a related technical discipline.
- Programming: Strong Python expertise is mandatory, plus proficiency in Java, Go, or C++.
- Production systems: Experience building and operating software used by real customers at scale, including deployment pipelines, monitoring, logging, observability, distributed tracing, high availability, fault tolerance, scalability, and disaster recovery.
- Security: Practical experience identifying vulnerabilities, implementing secure fixes, and validating remediation. Familiarity with injection, authorization, memory safety, race conditions, deserialization, secrets management, and input validation issues is required.
How it works
Apply on OpenTrain with your resume and then complete the application on the hiring site.
About AI training work
AI training work uses human-created examples, reviews, and evaluations to improve how artificial intelligence systems behave. This role focuses on testing coding agents, and experienced software engineers are needed to judge whether their solutions are useful, secure, and aligned with developer intent.