You will create and evaluate software engineering tasks that measure whether coding agents can complete work while preserving safety, privacy, permissions, data integrity, and developer intent. The work uses production-like repositories, runtime scenarios, configurations, documentation, tests, evaluators, reference patches, and scoring rubrics.
- Build benchmark tasks for bug fixes, feature additions, CI repairs, configuration migrations, integration updates, and runtime-state fixes.
- Define utility and safety requirements, then create visible and hidden tests for each task.
- Develop safe and intentionally unsafe reference solutions for evaluation scenarios.
- Assess shortcuts such as disabling tests, weakening assertions, deleting protected data, leaking secrets, broadening permissions, or bypassing validation.
- Package prompts, metadata, runnable repositories, evaluators, reference patches, rubrics, and calibration notes.
- Analyze coding-agent rollouts for utility and safety violations.
- Work with engineering, quality assurance, and security stakeholders to improve evaluator reliability and benchmark difficulty.
What it pays and takes
Pay details are not provided in the listing. The role is structured as part-time contract work and requires at least 20 hours each week.
- Open worldwide.
- English fluency required.
- At least eight years of hands-on software engineering experience in leading product companies or technology startups.
- Bachelor's or master's degree in computer science or a related technical field.
- Mandatory Python expertise, plus proficiency in Java, Go, or C++.
- Experience building and maintaining production software used by real customers.
- Strong skills in software architecture, debugging, performance optimization, deployment pipelines, monitoring, logging, and distributed tracing.
- Knowledge of high availability, fault tolerance, scalability, disaster recovery, secure code review, vulnerability remediation, and complex production failure analysis.
How it works
Apply on OpenTrain with your resume, then complete the application on the hiring site.
About AI training work
AI training work is the human work behind modern artificial intelligence, including preparing examples, evaluating model behavior, and testing whether systems follow instructions. Experienced software engineers are needed for specialist projects such as coding-agent evaluation, where technical judgment helps measure quality, reliability, and safe behavior.