Design realistic coding-agent benchmarks that test software engineering performance, safety, and security. Use your production engineering expertise on flexible, remote AI training work with OpenTrain.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Aug 6, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help contributors discover specialized projects, build a professional profile, and grow their experience in a rapidly expanding field where people directly shape how AI systems work.
About AI Training Work
AI training is the human side of building modern artificial intelligence. Experts create, review, and evaluate examples that help AI models and coding agents produce more useful, reliable, and responsible results.
Work remotely with a computer and internet connection
Contribute to cutting-edge AI systems through expert evaluation
Build experience that can become part of a long-term AI training portfolio
The Role
OpenTrain is hiring a Senior Coding Benchmark Software Engineer to create and evaluate realistic tasks for coding agents. You will work with production-like repositories, tests, configurations, documentation, and runtime scenarios to build benchmarks that measure engineering task completion and adherence to safety, security, privacy, permissions, and developer intent.
This contractor, part-time opportunity requires 20 or more hours per week and is open worldwide. English is required. The listing identifies the opportunity as entry level, while the role requirements call for at least eight years of hands-on software engineering experience.
What You'll Do
You will design benchmark tasks and develop the supporting materials needed to evaluate coding-agent behavior consistently. You will also analyze agent rollouts and work with engineering, quality, security, and client stakeholders to improve evaluator reliability and benchmark difficulty.
Design tasks covering bug fixes, feature additions, CI repairs, configuration migrations, integration updates, and runtime-state fixes
Define utility and safety requirements for each benchmark task
Build visible and hidden test suites that assess task completion and unsafe behavior
Analyze coding-agent rollouts and improve benchmark quality and difficulty
Required Qualifications
This role requires senior-level production software engineering experience and a strong understanding of reliable, secure, and observable systems.
At least eight years of hands-on software engineering experience
Experience at leading product companies or technology startups
Bachelor's or master's degree in computer science or a related technical discipline
Mandatory Python expertise plus experience with Java, Go, or C++
Experience building and maintaining production software used by real customers
Strong knowledge of software architecture, debugging, performance optimization, deployment pipelines, monitoring, logging, and observability
Experience diagnosing production failures using logs, metrics, and distributed tracing
Knowledge of high availability, fault tolerance, scalability, disaster recovery, and incident response
Secure code review experience, including common vulnerabilities, secure remediation, and validation of security fixes
Ability to design coding-agent benchmarks with utility and safety requirements
Knowledge of resilient, observable, large-scale production systems
Why Work With OpenTrain
OpenTrain gives contributors a focused way to build experience in AI training and data labeling while developing a credible professional portfolio. Your work on coding-agent evaluation can help demonstrate specialized expertise across future opportunities in this fast-growing industry.
Flexible part-time contractor work
Worldwide remote opportunity
Work directly on coding-agent evaluation and AI reliability
Build a profile that showcases your AI training experience
Evaluate and improve AI-generated code while creating benchmarking datasets for advanced software engineering models. This expert, remote contract role is part time at less than 20 hours per week.