Skip to content
OpenTrain AIFor AI Companies

ML Researcher Evaluation Specialist

Join OpenTrain as an ML Researcher Evaluation Specialist to assess frontier machine-learning researchers and their publication records for a high-priority pilot; part-time contractor role (20+ hrs/week), remote worldwide, paid $60–$100 USD/hr.

OpenTrain AI

Coding & Software

100% Remote Hourly · $60–$100/hr

$60–$100/hr

Compensation

Worldwide

Eligibility

Entry

Experience

Jul 29, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for building careers in AI training and data labeling. We make it simple to find specialized projects, build a unified portfolio, and grow a durable freelance career in a fast-growing industry.

As the hiring organization for this project, OpenTrain runs the engagement directly and supports contractors throughout the contract lifecycle.

About AI training and this pilot

AI training work is the human side of building modern models: people annotate, evaluate, and curate the examples and judgments that teach AI systems how to behave. This project is a short pilot focused on identifying top-tier research talent that drives algorithmic innovation.

The pilot emphasizes frontier ML topics—reinforcement learning, meta-learning, recursive self-improvement, and AI for science—and requires fast, high-quality assessments of research originality and impact.

The role

You will be the subject-matter expert who reviews candidate dossiers and publication records to identify researchers with original contributions in frontier ML. This is a part-time contractor role, remote and open worldwide, with an expected commitment of 20+ hours per week and an urgent turnaround cadence.

Work product will be documented evaluations (document-type inputs) and evaluation ratings that help prioritize strong research candidates for follow-up.

What you'll do

  • Review candidate backgrounds for originality and impact in frontier ML research.
  • Evaluate publication records at top-tier conferences (ICML, NeurIPS, ICLR) and judge significance of main conference papers.
  • Identify researchers aligned with reinforcement learning, meta-learning, recursive self-improvement, and AI for science.
  • Produce clear, timely evaluation ratings and written notes; prioritize speed without sacrificing quality.

Requirements

  • PhD in Machine Learning, Computer Science, AI, or a closely related field (required).
  • At least one main conference paper at ICML, NeurIPS, or ICLR (required).
  • Demonstrated experience conducting original ML research—ability to assess technical novelty and impact.
  • Preferred: multiple publications at top-tier ML conferences and expertise in reinforcement learning, meta-learning, recursive self-improvement, or AI for science.
  • Fluent English (project language) and ability to work 20+ hours per week as a contractor.

Compensation, schedule, and employment type

This is a contract, part-time engagement. OpenTrain hires contractors directly for this pilot.

Compensation: USD $60–$100 per hour (range provided by the project). Work is remote and open worldwide.

How it works and how to apply

Apply via your OpenTrain profile: submit your CV, list of publications, and brief notes on your research areas. You will review document-type candidate materials and record evaluation ratings using the platform’s interface.

The work is fast-paced and requires clear judgments and written rationale. Successful applicants will be asked to start quickly for the duration of the pilot.

  • Work involves reviewing documents and submitting evaluation ratings (label type: EVALUATION_RATING).
  • OpenTrain supports contractors with onboarding materials and the project rubric—bring deep ML research judgment and readiness to move quickly.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

LLM Evaluation Software Engineer (Ruby)

Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.

Coding & Software
Text
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026

C++ LLM Evaluation Software Engineer

Join OpenTrain as a remote contractor building and evaluating LLM performance on real C++ codebases; flexible 20/30/40 hr/week schedules and opportunities to lead junior engineers. Work with researchers to design verifiable engineering tasks, triage issues, run code, and rate model outputs.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026

Senior LLM Code Evaluation Engineer

Join OpenTrain AI to build evaluation datasets from public open-source code and measure how LLMs handle real-world software tasks; requires 3+ years software engineering with strong Go skills, 20+ hours/week, and eligibility in select countries.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026