Design and author multi-step scientific evaluation tasks for frontier AI models in a full-time remote US contractor role paying $60–$90/hr. Expect ~35 hours/week building Python reference solutions, defining rigorous criteria, and reviewing model attempts.
Generative AI & RLHF
Remote Hourly · $60–$90/hr
$60–$90/hr
Compensation
1 country
Eligibility
Entry
Experience
Jul 29, 2026
Posted
Open to applicants in
United States
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for building careers in AI training and data labeling. We make it easy to discover projects, manage your work, and grow a professional portfolio that showcases your contributions to how AI systems are built.
As the hiring organization for this role, OpenTrain coordinates the project, payments, and collaboration so you can focus on designing rigorous evaluation work that shapes real AI behavior.
Why AI training work matters
AI training (data labeling and evaluation) is the human foundation of modern machine learning—people annotate, test, and rate model outputs so models learn useful, reliable behaviors.
This type of work is remote, flexible, and directly influences how state-of-the-art systems reason, code, and analyze data. Contributors take on tasks that require domain knowledge, creativity, and careful judgment.
Role overview
We are hiring an AI Evaluation Benchmark Researcher to design and author complex, multi-step benchmarks that assess model reasoning, coding, and data-analysis skills. You will create tasks, write Python reference solutions and notebooks, and establish clear criteria that distinguish sound scientific reasoning from plausible-sounding errors.
You will work closely with research collaborators and other experts to align evaluations, review model attempts, and flag mistakes with the precision expected of an active researcher.
What you'll do
Design multi-step evaluation tasks that test experimental design, hypothesis formulation, analysis, and interpretation.
Author reference solutions in Python and Jupyter/Colab notebook formats that demonstrate expected reasoning and analysis.
Define detailed grading criteria to separate correct scientific reasoning from convincing but incorrect answers.
Review and rate model attempts on your tasks, marking errors and offering precise diagnostic feedback.
Coordinate with research collaborators and fellow evaluators to ensure consistency across benchmarks and tasks.
Requirements
You must meet all required qualifications listed below; this role expects careful, research-grade task design and evaluation.
MSc or PhD in a STEM field, or a computationally intensive social science or humanities discipline.
Minimum 1 year of active research experience in academia, industry, or a national lab.
Proven computational work using Python for analysis, simulation, modeling, or data pipelines.
Strong grounding in experimental design, hypothesis testing, and rigorous evaluation of results.
Working familiarity with Git, IDEs, and Jupyter or Colab notebook environments.
Excellent attention to detail, creativity in task design, and strong written communication.
Ability to commit approximately 35 hours per week and be available as a remote contractor based in the United States.
Eligible to work in the United States and able to receive weekly compensation.
Helpful background and who should apply
Prior experience authoring evaluation tasks, benchmarks, or working on model evaluation projects is a plus but not required.
This role fits researchers who enjoy turning scientific methodology into measurable tasks, writing reproducible Python analyses, and giving careful, research-standard feedback on model outputs.
Compensation, schedule, and how it works
This is a remote contractor role based in the United States. Compensation is $60–$90 per hour, paid weekly. The expected commitment is approximately 35 hours per week.
Employment type: contractor / part-time. You will work through OpenTrain, which manages contracts, payments, and project coordination so you can focus on benchmark design and evaluation.
Pay: USD $60–$90 per hour, paid weekly.
Work location: Remote — must be eligible to work in the United States.
Schedule: Approximately 35 hours per week (flexible within the expectation of consistent weekly availability).
Review AI-generated market research outputs and produce gold-standard briefs, surveys, and insight syntheses in a remote, hourly contractor role. Part-time (20+ hrs/week), US$30–65/hr; requires 5+ years in market research and C1 English.
Join OpenTrain AI as a Machine Learning Engineer to design and run evaluation benchmarks for frontier LLMs. Contract, part-time role (20+ hrs/week) paying $60–$90/hr for US-based practitioners with hands-on ML training and research experience.
Assess AI-generated reproductions of published ML papers: read papers, inspect code/data/output artifacts, rank reproduction attempts, and provide clear written evaluations. Remote, contractor role for experienced ML researchers — 20+ hrs/week.