Create and evaluate realistic software-engineering benchmarks for coding agents using production-like repositories, secure coding, and rigorous testing. This fully remote contractor assignment runs 4 to 8 weeks.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Sep 2, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover cutting-edge projects, build a credible professional profile, and apply in minutes. Creating an OpenTrain account is free.
OpenTrain AI is recruiting for this contractor assignment. Your work will contribute to the human side of artificial intelligence, where experienced professionals help evaluate and improve the systems behind modern AI.
About AI Training and Coding-Agent Evaluation
AI training includes evaluating model outputs, designing challenging examples, and reviewing whether AI systems follow instructions safely and accurately. Coding-agent benchmarks apply this work to software engineering by testing how well agents modify repositories, resolve failures, and respect technical and security constraints.
This is part of a fast-growing field in which human expertise directly shapes how state-of-the-art AI systems behave. The work is remote and can offer flexible, project-based opportunities for specialists.
The Role
As a Senior Coding-Agent Benchmark Engineer, you will create and evaluate realistic software-engineering tasks used to assess coding agents. You will work with production-like repositories, tests, configurations, documentation, and runtime scenarios, translating practical engineering problems into rigorous benchmark environments.
The role combines software development, production debugging, secure coding, and AI model evaluation. You will assess both successful task completion and unsafe behavior, including shortcuts that appear effective but violate system or security requirements.
Fully remote contractor assignment
Expected commitment of at least 4 hours per day and 20 hours per week
Includes 4 hours of overlap with Pacific Time
Assignment duration of 4 to 8 weeks
Compensation is not disclosed
What You'll Do
You will design complete benchmark environments and the materials needed to run, evaluate, and calibrate them. Your work will require strong engineering judgment, careful attention to system constraints, and the ability to distinguish legitimate solutions from unsafe shortcuts.
Design tasks involving bug fixes, feature additions, CI repairs, configuration migrations, integration updates, and runtime-state fixes.
Define utility and safety requirements that verify completion while preserving system constraints, developer intent, data integrity, privacy, permissions, and oversight mechanisms.
Develop visible and hidden test suites, safe and unsafe reference solutions, evaluators, scoring rubrics, and calibration materials.
Analyze coding-agent rollouts for completion and unsafe behavior, including disabling tests, weakening assertions, deleting protected data, leaking secrets, broadening permissions, or bypassing validation.
Package prompts, metadata, runnable repositories, evaluators, reference patches, and related benchmark materials.
Collaborate with engineering, QA, security, and client teams to improve evaluator reliability and benchmark difficulty.
Required Qualifications
This role is intended for an experienced software engineer with a strong production background and the ability to reason across implementation, operations, reliability, and security. The role requires a bachelor's or master's degree in computer science or a related technical discipline and at least eight years of hands-on software engineering experience in product companies or technology startups.
Mandatory Python programming expertise plus working knowledge of Java, Go, or C++.
Experience building and maintaining production software at scale.
Experience with deployment pipelines, monitoring, logging, and distributed tracing.
Ability to diagnose complex production failures.
Knowledge of architecture, performance optimization, high availability, fault tolerance, scalability, and disaster recovery.
Knowledge of secure software development, including identifying vulnerabilities and validating remediation.
Security and Evaluation Judgment
A central part of the role is determining whether a coding agent solved a task correctly without undermining the surrounding system. You should be comfortable reviewing implementation choices for both functional quality and safety.
Identify vulnerabilities such as injection, authorization flaws, race conditions, deserialization issues, and input validation failures.
Implement secure fixes and validate that remediation is effective.
Recognize shortcuts that weaken tests, permissions, privacy, or data integrity.
Distinguish legitimate coding-agent task completion from behavior that bypasses validation or oversight.
Who Should Apply
Apply if you are a senior production software engineer who enjoys turning real engineering challenges into precise, reproducible evaluations. This opportunity is especially suited to people who can move between code, infrastructure, reliability, security, and technical assessment without losing sight of developer intent or system constraints.
The structured experience field identifies this project as entry level, but the role description requires the senior qualifications and professional experience listed above. Review those requirements carefully before applying.
Experienced production software engineers
Engineers familiar with distributed systems and operational debugging
Security-minded developers who can assess vulnerabilities and remediation
Technical specialists interested in shaping the behavior of coding agents
How to Work With OpenTrain
OpenTrain helps contributors discover AI training opportunities, build a profile around credible experience, and grow toward a durable career in this rapidly expanding field. Create a free account to manage your profile and apply for suitable projects in minutes.
Review the role requirements and schedule.
Prepare evidence of relevant software engineering, production, and security experience.
Apply through OpenTrain for consideration.
Complete the assignment remotely during the stated engagement period.
Build realistic coding-agent benchmarks, test suites, and security-focused evaluators for large language models. This part-time contractor role is open worldwide and requires senior production engineering experience.
Review and benchmark AI-generated software as a remote contractor in the United States. Use your software engineering expertise to test code, assess explanations, investigate failures, and improve coding evaluation standards.
Use your Python and AI engineering experience to evaluate coding-agent trajectories, tool calls, code changes, and technical outcomes. Work worldwide on a flexible 20+ hour-per-week contract through OpenTrain.