Skip to content
OpenTrain AIFor AI Companies

AI Code Evaluation and Benchmarking Engineer

Evaluate and benchmark AI-generated code: review correctness, debug and verify solutions, and build evaluation datasets for frontier models. US-remote, contractor role — 20+ hrs/week (min 4 hrs/day), 1-month contract with 4-hour PST overlap required.

OpenTrain AI

Coding & Software

Remote

1 country

Eligibility

Entry

Experience

Jul 17, 2026

Posted

Open to applicants in

United States

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the leading platform for people building careers in AI training and data labeling. We connect skilled contributors with meaningful evaluation and annotation work, enabling you to build a portfolio, find flexible remote projects, and grow into a durable freelance career in a fast-growing industry.

For this role, OpenTrain AI is the hiring and contracting organization. You will work directly with our project teams to evaluate model outputs, improve benchmarks, and help shape how AI systems are measured and improved.

About AI Training Work

AI training (also called data labeling or human feedback work) is the human side of building modern AI: people prepare, review, and grade examples that models learn from. Contributors do tasks like annotating text, images, or code; rating and ranking model outputs; and refining benchmarks that guide model development.

This work is 100% remote, often flexible and part-time, and accessible to many contributors—while specialist tasks such as code evaluation pay more for software engineering experience. By joining OpenTrain you help shape cutting-edge models while keeping flexible hours.

The Role

We are recruiting an AI Code Evaluation Engineer to assess and benchmark the coding capabilities of frontier AI models. You will evaluate AI-generated code for correctness, quality, and adherence to requirements, reproduce and debug issues, and help create high-quality evaluation datasets and rubrics.

This position suits engineers who enjoy code review, debugging, problem-solving, and applying strong software engineering judgment to technical tasks.

What You'll Do

  • Review and evaluate AI-generated code for correctness, efficiency, maintainability, and adherence to requirements.
  • Analyze software engineering tasks and validate whether proposed solutions meet expected outcomes.
  • Debug code, reproduce issues, and verify fixes across different programming environments.
  • Assess model-generated explanations, reasoning, and implementation approaches for technical accuracy.
  • Create, refine, and maintain evaluation datasets, benchmarks, and grading rubrics for coding tasks.
  • Identify edge cases, failure modes, and areas where AI systems struggle with software engineering problems.
  • Document findings clearly and provide structured feedback to improve evaluation quality and consistency.
  • Collaborate with project teams to establish quality standards and evaluation methodologies.

Requirements

You must meet all of the stated requirements below to be considered for this role.

  • Bachelor's or Master's degree in Computer Science, Software Engineering, or a related technical field.
  • 3+ years of professional software engineering experience.
  • Strong proficiency in one or more of: Python, Java, C/C++, Go, Swift, Objective-C, PHP, or SQL.
  • Strong understanding of data structures, algorithms, software design principles, and debugging methodologies.
  • Experience performing code reviews and evaluating code quality in production or large-scale codebases.
  • Familiarity with version control systems such as Git and modern software development workflows.
  • Ability to analyze complex technical problems and assess solution correctness with minimal supervision.
  • Strong written communication skills and attention to detail.

Helpful Background

The following experiences are not required but will help you excel in this role:

  • Experience with AI/ML data annotation, NLP, prompt engineering, model evaluation, or LLM-related projects.
  • Experience evaluating AI-generated code, benchmark creation, or software quality assessment.

Schedule, Location, and Contract Details

This is a remote contractor role available to applicants located in the United States only.

Time commitment: at least 4 hours per day with a minimum of 20 hours per week. You must have 4 hours of overlap with Pacific Standard Time (PST). Contract duration: 1 month. Employment types: contractor, part-time.

  • Location: US only (remote).
  • Minimum weekly commitment: 20 hours (at least 4 hours/day).
  • PST overlap: 4 hours required.
  • Contract length: 1 month.

Who Should Apply and How It Works

Apply if you are an experienced software engineer who enjoys reviewing code, diagnosing tricky bugs, and formalizing evaluation criteria. You will help determine how well AI systems perform on real engineering tasks and contribute to benchmarks that guide model improvement.

OpenTrain makes it easy to build an AI training portfolio and find flexible remote projects. If this role fits your skills and availability, create or update your OpenTrain profile and apply — our team will review your qualifications and follow up with next steps.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Benchmark Engineer — Software Engineering

Design and validate multi-agent coding benchmarks using real open-source code changes, Docker, and Python verification scripts. Remote 4-week contractor role for developers in select countries, requiring daily overlap with PST.

Coding & Software
Computer Code Programming
Remote · Bangladesh, Brazil, Colombia +8 more
English
Part-time · Flexible
Entry level

Posted Jul 24, 2026

Machine Learning Engineer, Benchmarking & Evaluation

Join OpenTrain as a remote Machine Learning Engineer focused on benchmark-driven evaluation of real-world ML systems. This contractor role requires 3+ years of ML engineering experience, strong Python skills, and availability 20+ hrs/week with PST overlap.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +7 more
English
Part-time · Flexible
Intermediate level

Posted Jul 16, 2026

Machine Learning Benchmark Evaluator

Experienced ML engineers wanted for part-time, remote contract work evaluating production-grade model training, evaluation, and inference pipelines; requires 3+ years of ML engineering experience, strong Python, and PyTorch/TensorFlow/JAX familiarity. 20+ hours/week; apply through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +7 more
English
Part-time · Flexible
Expert level

Posted Jul 17, 2026