Join OpenTrain as an ML Researcher Evaluation Specialist to assess frontier machine-learning researchers and their publication records for a high-priority pilot; part-time contractor role (20+ hrs/week), remote worldwide, paid $60–$100 USD/hr.
Coding & Software
100% Remote Hourly · $60–$100/hr
$60–$100/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 29, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for building careers in AI training and data labeling. We make it simple to find specialized projects, build a unified portfolio, and grow a durable freelance career in a fast-growing industry.
As the hiring organization for this project, OpenTrain runs the engagement directly and supports contractors throughout the contract lifecycle.
About AI training and this pilot
AI training work is the human side of building modern models: people annotate, evaluate, and curate the examples and judgments that teach AI systems how to behave. This project is a short pilot focused on identifying top-tier research talent that drives algorithmic innovation.
The pilot emphasizes frontier ML topics—reinforcement learning, meta-learning, recursive self-improvement, and AI for science—and requires fast, high-quality assessments of research originality and impact.
The role
You will be the subject-matter expert who reviews candidate dossiers and publication records to identify researchers with original contributions in frontier ML. This is a part-time contractor role, remote and open worldwide, with an expected commitment of 20+ hours per week and an urgent turnaround cadence.
Work product will be documented evaluations (document-type inputs) and evaluation ratings that help prioritize strong research candidates for follow-up.
What you'll do
Review candidate backgrounds for originality and impact in frontier ML research.
Evaluate publication records at top-tier conferences (ICML, NeurIPS, ICLR) and judge significance of main conference papers.
Identify researchers aligned with reinforcement learning, meta-learning, recursive self-improvement, and AI for science.
Produce clear, timely evaluation ratings and written notes; prioritize speed without sacrificing quality.
Requirements
PhD in Machine Learning, Computer Science, AI, or a closely related field (required).
At least one main conference paper at ICML, NeurIPS, or ICLR (required).
Demonstrated experience conducting original ML research—ability to assess technical novelty and impact.
Preferred: multiple publications at top-tier ML conferences and expertise in reinforcement learning, meta-learning, recursive self-improvement, or AI for science.
Fluent English (project language) and ability to work 20+ hours per week as a contractor.
Compensation, schedule, and employment type
This is a contract, part-time engagement. OpenTrain hires contractors directly for this pilot.
Compensation: USD $60–$100 per hour (range provided by the project). Work is remote and open worldwide.
How it works and how to apply
Apply via your OpenTrain profile: submit your CV, list of publications, and brief notes on your research areas. You will review document-type candidate materials and record evaluation ratings using the platform’s interface.
The work is fast-paced and requires clear judgments and written rationale. Successful applicants will be asked to start quickly for the duration of the pilot.
Work involves reviewing documents and submitting evaluation ratings (label type: EVALUATION_RATING).
OpenTrain supports contractors with onboarding materials and the project rubric—bring deep ML research judgment and readiness to move quickly.
Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.
Join OpenTrain as a remote contractor building and evaluating LLM performance on real C++ codebases; flexible 20/30/40 hr/week schedules and opportunities to lead junior engineers. Work with researchers to design verifiable engineering tasks, triage issues, run code, and rate model outputs.
Join OpenTrain AI to build evaluation datasets from public open-source code and measure how LLMs handle real-world software tasks; requires 3+ years software engineering with strong Go skills, 20+ hours/week, and eligibility in select countries.