Skip to content
OpenTrain AIFor AI Companies

Machine Learning Engineer — Model Evaluation & Experimentation

Join OpenTrain AI as a Machine Learning Engineer to design and run evaluation benchmarks for frontier LLMs. Contract, part-time role (20+ hrs/week) paying $60–$90/hr for US-based practitioners with hands-on ML training and research experience.

OpenTrain AI

Generative AI & RLHF

Remote Hourly · $60–$90/hr

$60–$90/hr

Compensation

1 country

Eligibility

Intermediate

Experience

Jul 25, 2026

Posted

Open to applicants in

United States

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover specialized projects, build a unified AI training portfolio, and grow a durable freelance career in an industry that shapes how modern AI systems learn.

About AI training and this work

AI training (data annotation, human feedback, and model evaluation) is the human side of creating reliable AI. Contributors do everything from designing benchmarks to rating outputs and running experiments that teach models what correct behavior looks like.

This role sits at the cutting edge: you will author and run multi-step evaluation experiments that directly inform how next-generation language models are measured and improved.

The role

OpenTrain AI is hiring an experienced machine learning practitioner to act as a ground-truth expert for frontier model evaluation. You will turn research ideas into reproducible evaluation tasks, run training experiments, and analyze results to show correct solutions and where models fall short.

This is a contract, part-time position for US-based contributors, expected at 20+ hours per week.

What you'll do

You will design, implement, and run end-to-end evaluation and experimentation workflows that benchmark language models and training approaches.

  • Turn vague ML research ideas into well-defined, multi-step evaluation tasks and benchmarks.
  • Implement changes, run training experiments, and analyze outcomes to demonstrate correct behavior.
  • Build tasks and analyses around reinforcement learning concepts such as reward functions and training behavior.
  • Evaluate frontier models on your tasks, document failure modes, and explain where and why they fall short.
  • Collaborate with researchers and other experts to keep tasks consistent, rigorous, and reproducible.

Requirements

You must meet the core qualifications and be able to work independently on open-ended research problems.

  • MSc or PhD in machine learning, computer science, or a related STEM field (or equivalent experience).
  • 1+ years in a research or research-engineering role.
  • Hands-on experience training and evaluating ML models end-to-end, including experiment setup, execution, and analysis.
  • Strong familiarity with large language models, their capabilities, limitations, and evaluation techniques.
  • Working proficiency in Python and Git, comfortable in both scripting and notebook environments.
  • High attention to detail, creativity in task design, and strong written communication skills.
  • Ability to work independently on ambiguous, open-ended problems.
  • Must be legally able to work in the United States.

Helpful background and tools

The following are not required but will help you move faster in the role.

  • Basic understanding of reinforcement learning (reward functions, policy training).
  • Past experience in AI training, model evaluation, or benchmark/task authoring.
  • Familiarity with evaluation workflows for text tasks, including rating schemas and text-generation assessment.

Compensation, schedule, and how it works

This is a contractor, part-time role with flexible hours. Expect to contribute 20+ hours per week and manage your own schedule within agreed experiment timelines.

  • Pay: Hourly contract at $60–$90 USD per hour (posted rate: $90/hr; per-role range $60–$90/hr).
  • Work type: Remote, US-only applicants.
  • Data & tasks: Text-focused evaluation work including evaluation ratings and text-generation benchmark design.
  • Employment: Contract / Part-time; OpenTrain AI is the hiring organization and project owner.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Model Evaluation Developer

Join OpenTrain as an AI Model Evaluation Developer to write and maintain code, run model benchmarks, rank responses, and build datasets for fine-tuning and RLHF. This remote, part-time contractor role requires strong Python and JavaScript/TypeScript skills and 20+ hours/week.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Intermediate level

Posted Jul 17, 2026

ML & NLP Evaluation Expert (Remote US, 20 hrs/wk)

Join OpenTrain as a part-time ML & NLP Evaluation Expert designing robust evaluation tasks, reference solutions, and assessing frontier model outputs; remote within the US, ~20 hours/week, $80–$110/hr.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Entry level
Hourly · $80–$110/hr

Posted Jul 13, 2026

ML Research Paper Reproduction Evaluator

Assess AI-generated reproductions of published ML papers: read papers, inspect code/data/output artifacts, rank reproduction attempts, and provide clear written evaluations. Remote, contractor role for experienced ML researchers — 20+ hrs/week.

Generative AI & RLHF
Document
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 25, 2026