Skip to content
OpenTrain AIFor AI Companies

Senior Coding Benchmark Software Engineer

Design realistic coding-agent benchmarks that test software engineering performance, safety, and security. Use your production engineering expertise on flexible, remote AI training work with OpenTrain.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Entry

Experience

Aug 6, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help contributors discover specialized projects, build a professional profile, and grow their experience in a rapidly expanding field where people directly shape how AI systems work.

About AI Training Work

AI training is the human side of building modern artificial intelligence. Experts create, review, and evaluate examples that help AI models and coding agents produce more useful, reliable, and responsible results.

  • Work remotely with a computer and internet connection
  • Contribute to cutting-edge AI systems through expert evaluation
  • Build experience that can become part of a long-term AI training portfolio

The Role

OpenTrain is hiring a Senior Coding Benchmark Software Engineer to create and evaluate realistic tasks for coding agents. You will work with production-like repositories, tests, configurations, documentation, and runtime scenarios to build benchmarks that measure engineering task completion and adherence to safety, security, privacy, permissions, and developer intent.

This contractor, part-time opportunity requires 20 or more hours per week and is open worldwide. English is required. The listing identifies the opportunity as entry level, while the role requirements call for at least eight years of hands-on software engineering experience.

What You'll Do

You will design benchmark tasks and develop the supporting materials needed to evaluate coding-agent behavior consistently. You will also analyze agent rollouts and work with engineering, quality, security, and client stakeholders to improve evaluator reliability and benchmark difficulty.

  • Design tasks covering bug fixes, feature additions, CI repairs, configuration migrations, integration updates, and runtime-state fixes
  • Define utility and safety requirements for each benchmark task
  • Build visible and hidden test suites that assess task completion and unsafe behavior
  • Create safe and unsafe reference solutions
  • Package prompts, metadata, runnable repositories, evaluators, reference patches, scoring rubrics, and calibration notes
  • Analyze coding-agent rollouts and improve benchmark quality and difficulty

Required Qualifications

This role requires senior-level production software engineering experience and a strong understanding of reliable, secure, and observable systems.

  • At least eight years of hands-on software engineering experience
  • Experience at leading product companies or technology startups
  • Bachelor's or master's degree in computer science or a related technical discipline
  • Mandatory Python expertise plus experience with Java, Go, or C++
  • Experience building and maintaining production software used by real customers
  • Strong knowledge of software architecture, debugging, performance optimization, deployment pipelines, monitoring, logging, and observability
  • Experience diagnosing production failures using logs, metrics, and distributed tracing
  • Knowledge of high availability, fault tolerance, scalability, disaster recovery, and incident response
  • Secure code review experience, including common vulnerabilities, secure remediation, and validation of security fixes
  • Ability to design coding-agent benchmarks with utility and safety requirements
  • Knowledge of resilient, observable, large-scale production systems

Why Work With OpenTrain

OpenTrain gives contributors a focused way to build experience in AI training and data labeling while developing a credible professional portfolio. Your work on coding-agent evaluation can help demonstrate specialized expertise across future opportunities in this fast-growing industry.

  • Flexible part-time contractor work
  • Worldwide remote opportunity
  • Work directly on coding-agent evaluation and AI reliability
  • Build a profile that showcases your AI training experience

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Software Engineer, AI Coding Benchmark Development

OpenTrain AI is hiring a senior software engineer to design production-like coding benchmarks and safety tests for frontier coding agents. Contractor role: remote, 20+ hrs/week (min 4 hrs/day), 4–8 week assignment; Python required and 8+ years' experience.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Expert level

Posted Aug 1, 2026

Software Engineering Code Evaluator

Evaluate and improve AI-generated code while creating benchmarking datasets for advanced software engineering models. This expert, remote contract role is part time at less than 20 hours per week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Expert level

Posted Jul 16, 2026

AI Code Evaluation and Benchmarking Engineer

Evaluate and benchmark AI-generated code: review correctness, debug and verify solutions, and build evaluation datasets for frontier models. US-remote, contractor role — 20+ hrs/week (min 4 hrs/day), 1-month contract with 4-hour PST overlap required.

Coding & Software
Text
Remote · United States
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026