Skip to content
OpenTrain AIFor AI Companies

AI Evaluation Benchmark Researcher

Design and author multi-step scientific evaluation tasks for frontier AI models in a full-time remote US contractor role paying $60–$90/hr. Expect ~35 hours/week building Python reference solutions, defining rigorous criteria, and reviewing model attempts.

OpenTrain AI

Generative AI & RLHF

Remote Hourly · $60–$90/hr

$60–$90/hr

Compensation

1 country

Eligibility

Entry

Experience

Jul 29, 2026

Posted

Open to applicants in

United States

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for building careers in AI training and data labeling. We make it easy to discover projects, manage your work, and grow a professional portfolio that showcases your contributions to how AI systems are built.

As the hiring organization for this role, OpenTrain coordinates the project, payments, and collaboration so you can focus on designing rigorous evaluation work that shapes real AI behavior.

Why AI training work matters

AI training (data labeling and evaluation) is the human foundation of modern machine learning—people annotate, test, and rate model outputs so models learn useful, reliable behaviors.

This type of work is remote, flexible, and directly influences how state-of-the-art systems reason, code, and analyze data. Contributors take on tasks that require domain knowledge, creativity, and careful judgment.

Role overview

We are hiring an AI Evaluation Benchmark Researcher to design and author complex, multi-step benchmarks that assess model reasoning, coding, and data-analysis skills. You will create tasks, write Python reference solutions and notebooks, and establish clear criteria that distinguish sound scientific reasoning from plausible-sounding errors.

You will work closely with research collaborators and other experts to align evaluations, review model attempts, and flag mistakes with the precision expected of an active researcher.

What you'll do

  • Design multi-step evaluation tasks that test experimental design, hypothesis formulation, analysis, and interpretation.
  • Author reference solutions in Python and Jupyter/Colab notebook formats that demonstrate expected reasoning and analysis.
  • Define detailed grading criteria to separate correct scientific reasoning from convincing but incorrect answers.
  • Review and rate model attempts on your tasks, marking errors and offering precise diagnostic feedback.
  • Coordinate with research collaborators and fellow evaluators to ensure consistency across benchmarks and tasks.

Requirements

You must meet all required qualifications listed below; this role expects careful, research-grade task design and evaluation.

  • MSc or PhD in a STEM field, or a computationally intensive social science or humanities discipline.
  • Minimum 1 year of active research experience in academia, industry, or a national lab.
  • Proven computational work using Python for analysis, simulation, modeling, or data pipelines.
  • Strong grounding in experimental design, hypothesis testing, and rigorous evaluation of results.
  • Working familiarity with Git, IDEs, and Jupyter or Colab notebook environments.
  • Excellent attention to detail, creativity in task design, and strong written communication.
  • Ability to commit approximately 35 hours per week and be available as a remote contractor based in the United States.
  • Eligible to work in the United States and able to receive weekly compensation.

Helpful background and who should apply

Prior experience authoring evaluation tasks, benchmarks, or working on model evaluation projects is a plus but not required.

This role fits researchers who enjoy turning scientific methodology into measurable tasks, writing reproducible Python analyses, and giving careful, research-standard feedback on model outputs.

Compensation, schedule, and how it works

This is a remote contractor role based in the United States. Compensation is $60–$90 per hour, paid weekly. The expected commitment is approximately 35 hours per week.

Employment type: contractor / part-time. You will work through OpenTrain, which manages contracts, payments, and project coordination so you can focus on benchmark design and evaluation.

  • Pay: USD $60–$90 per hour, paid weekly.
  • Work location: Remote — must be eligible to work in the United States.
  • Schedule: Approximately 35 hours per week (flexible within the expectation of consistent weekly availability).

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Market Research AI Evaluation Expert

Review AI-generated market research outputs and produce gold-standard briefs, surveys, and insight syntheses in a remote, hourly contractor role. Part-time (20+ hrs/week), US$30–65/hr; requires 5+ years in market research and C1 English.

Generative AI & RLHF
Text
Remote · Bangladesh, Bhutan, Brazil +14 more
English
Part-time · Flexible
Intermediate level
Hourly · $30–$65/hr

Posted Jul 9, 2026

Machine Learning Engineer, Model Evaluation & Experimentation

Join OpenTrain AI as a Machine Learning Engineer to design and run evaluation benchmarks for frontier LLMs. Contract, part-time role (20+ hrs/week) paying $60–$90/hr for US-based practitioners with hands-on ML training and research experience.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Intermediate level
Hourly · $60–$90/hr

Posted Jul 25, 2026

ML Research Paper Reproduction Evaluator

Assess AI-generated reproductions of published ML papers: read papers, inspect code/data/output artifacts, rank reproduction attempts, and provide clear written evaluations. Remote, contractor role for experienced ML researchers — 20+ hrs/week.

Generative AI & RLHF
Document
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 25, 2026