Machine Learning Engineer — Model Evaluation & Experimentation
Join OpenTrain AI as a Machine Learning Engineer to design and run evaluation benchmarks for frontier LLMs. Contract, part-time role (20+ hrs/week) paying $60–$90/hr for US-based practitioners with hands-on ML training and research experience.
Generative AI & RLHF
Remote Hourly · $60–$90/hr
$60–$90/hr
Compensation
1 country
Eligibility
Intermediate
Experience
Jul 25, 2026
Posted
Open to applicants in
United States
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover specialized projects, build a unified AI training portfolio, and grow a durable freelance career in an industry that shapes how modern AI systems learn.
About AI training and this work
AI training (data annotation, human feedback, and model evaluation) is the human side of creating reliable AI. Contributors do everything from designing benchmarks to rating outputs and running experiments that teach models what correct behavior looks like.
This role sits at the cutting edge: you will author and run multi-step evaluation experiments that directly inform how next-generation language models are measured and improved.
The role
OpenTrain AI is hiring an experienced machine learning practitioner to act as a ground-truth expert for frontier model evaluation. You will turn research ideas into reproducible evaluation tasks, run training experiments, and analyze results to show correct solutions and where models fall short.
This is a contract, part-time position for US-based contributors, expected at 20+ hours per week.
What you'll do
You will design, implement, and run end-to-end evaluation and experimentation workflows that benchmark language models and training approaches.
Turn vague ML research ideas into well-defined, multi-step evaluation tasks and benchmarks.
Implement changes, run training experiments, and analyze outcomes to demonstrate correct behavior.
Build tasks and analyses around reinforcement learning concepts such as reward functions and training behavior.
Evaluate frontier models on your tasks, document failure modes, and explain where and why they fall short.
Collaborate with researchers and other experts to keep tasks consistent, rigorous, and reproducible.
Requirements
You must meet the core qualifications and be able to work independently on open-ended research problems.
MSc or PhD in machine learning, computer science, or a related STEM field (or equivalent experience).
1+ years in a research or research-engineering role.
Hands-on experience training and evaluating ML models end-to-end, including experiment setup, execution, and analysis.
Strong familiarity with large language models, their capabilities, limitations, and evaluation techniques.
Working proficiency in Python and Git, comfortable in both scripting and notebook environments.
High attention to detail, creativity in task design, and strong written communication skills.
Ability to work independently on ambiguous, open-ended problems.
Must be legally able to work in the United States.
Helpful background and tools
The following are not required but will help you move faster in the role.
Basic understanding of reinforcement learning (reward functions, policy training).
Past experience in AI training, model evaluation, or benchmark/task authoring.
Familiarity with evaluation workflows for text tasks, including rating schemas and text-generation assessment.
Compensation, schedule, and how it works
This is a contractor, part-time role with flexible hours. Expect to contribute 20+ hours per week and manage your own schedule within agreed experiment timelines.
Pay: Hourly contract at $60–$90 USD per hour (posted rate: $90/hr; per-role range $60–$90/hr).
Work type: Remote, US-only applicants.
Data & tasks: Text-focused evaluation work including evaluation ratings and text-generation benchmark design.
Employment: Contract / Part-time; OpenTrain AI is the hiring organization and project owner.
Join OpenTrain as an AI Model Evaluation Developer to write and maintain code, run model benchmarks, rank responses, and build datasets for fine-tuning and RLHF. This remote, part-time contractor role requires strong Python and JavaScript/TypeScript skills and 20+ hours/week.
Join OpenTrain as a part-time ML & NLP Evaluation Expert designing robust evaluation tasks, reference solutions, and assessing frontier model outputs; remote within the US, ~20 hours/week, $80–$110/hr.
Assess AI-generated reproductions of published ML papers: read papers, inspect code/data/output artifacts, rank reproduction attempts, and provide clear written evaluations. Remote, contractor role for experienced ML researchers — 20+ hrs/week.