Join OpenTrain as a part-time ML & NLP Evaluation Expert designing robust evaluation tasks, reference solutions, and assessing frontier model outputs; remote within the US, ~20 hours/week, $80–$110/hr.
Generative AI & RLHF
Remote Hourly · $80–$110/hr
$80–$110/hr
Compensation
1 country
Eligibility
Entry
Experience
Jul 13, 2026
Posted
Open to applicants in
United States
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the centralized platform where people build careers in AI training and data labeling. We help freelancers find specialized, remote projects, consolidate opportunities, and grow a durable freelance career working directly on how AI systems are trained and evaluated.
We hire and contract experts directly for many projects. This role will be engaged, paid, and managed by OpenTrain.
Free to join and build a unified AI training portfolio on the OpenTrain platform.
Work remotely with flexible hours on projects that shape real AI systems.
About AI training and this work
AI training (data labeling, annotation, and evaluation) is the human side of building modern models: people design tasks, produce reference answers, and judge model outputs so researchers and engineers can measure and improve capabilities.
This position places you on the cutting edge of model evaluation—designing challenging, real-world tasks that reveal capability gaps in frontier models.
Work directly influences how state-of-the-art models are measured and improved.
Flexible, remote, part-time work that fits around other commitments.
The role
OpenTrain is hiring a Machine Learning & NLP Evaluation Expert to design evaluation tasks, craft reference solutions, and evaluate advanced model outputs. This is a part-time contractor role based in the United States, approximately 20 hours per week.
You will join a distributed team of experts working on a sustained evaluation effort focused on frontier ML/NLP models.
Hours: ~20 hours per week (part-time contractor).
Location: Remote, United States only.
Employment type: Contractor, part-time.
What you'll do
You will create robust, difficult evaluation tasks that target specific capability gaps in modern ML/NLP models, produce golden reference solutions and specs, integrate tasks into development environments, and assess model behaviour to identify failure modes.
Design and develop ML/NLP evaluation tasks that reflect real-world challenge cases.
Integrate tasks into an agentic development environment using Python.
Generate golden reference solutions and clear evaluation specifications.
Run evaluations of target models and document failure modes and capability gaps.
Collaborate with other subject-matter experts to maintain consistency and accuracy.
Requirements
You must have hands-on ML and NLP experience, Python proficiency, and a strong grasp of modern transformer-based methods and evaluation pipelines. Reliable availability and strong written communication are required.
Deep hands-on experience in machine learning and natural language processing from applied industry, research, or graduate/PhD work.
Working proficiency in Python applied in research, industry, or open-source projects.
Strong command of modern ML/NLP methods including transformer models, training and evaluation pipelines.
Reliable availability for approximately 20 hours per week and ability to work independently.
Excellent written communication and collaboration skills.
Helpful background and who should apply
Prior experience in AI training, model evaluation, or annotation is helpful but not strictly required if you meet the core ML/NLP and Python criteria. This role suits researchers, engineers, and advanced graduate students who enjoy designing adversarial and diagnostic tasks for models.
Researchers or engineers with experience creating evaluation benchmarks or challenge sets.
Graduate students with ML/NLP thesis work and strong Python tooling skills.
Practitioners interested in part-time, impactful work shaping model behaviour.
Compensation, data, and how to apply
Pay is $80–$110 per hour. This position focuses on text data and evaluation ratings (EVALUATION_RATING). You will be contracted and paid by OpenTrain.
To apply, create or use your OpenTrain account and submit your profile and application. Include examples of past ML/NLP work, evaluation tasks or benchmarks you’ve designed, and relevant code or publications if available.
Pay: $80–$110 per hour (hourly contractor).
Data type: Text; label type: Evaluation rating.
Language: English; work eligibility: United States only.
Create multi-turn conversations, rubrics, and evaluation assets for frontier LLMs while working remotely as a contractor 20+ hours/week. Rapid onboarding and clear specs; paid on a per-task/hour basis at $20–$30/hr.
Join OpenTrain to evaluate large language model outputs, create challenging prompts, and deliver recorded verbal feedback; remote, contract role (20+ hrs/week) paying $20–$30/hr. Entry-level friendly for strong American English speakers with LLM experience.
Join OpenTrain as a remote LLM Red-Teamer to design adversarial multi-turn conversations, write rigorous evaluation rubrics, and validate frontier language models 20+ hrs/week for $40–$65/hr. Prior RLHF or evaluation experience is helpful but not required.