ML Research Paper Reproduction Evaluator
Assess AI-generated reproductions of published ML papers: read papers, inspect code/data/output artifacts, rank reproduction attempts, and provide clear written evaluations. Remote, contractor role for experienced ML researchers — 20+ hrs/week.
Generative AI & RLHF
Worldwide
Eligibility
Entry
Experience
Jul 25, 2026
Posted
Open worldwide
About OpenTrain
OpenTrain is the centralized platform where people start and grow careers in AI training and data labeling. We connect skilled contributors with specialized, remote work that directly shapes how modern AI systems behave, and help freelancers build a unified portfolio they control.
- Work 100% remotely and apply flexibly to projects that fit your schedule.
- Build a durable freelance career in a fast-growing area of tech focused on human-in-the-loop AI.
About AI training and this role
AI training (data labeling and human feedback) is the human side of building AI: people create, evaluate, and improve the examples models learn from. This role focuses on evaluating AI-generated reproductions of published machine learning research to ensure technical correctness and methodological soundness.
As an ML Research Paper Reproduction Evaluator you'll use your research judgment to read original papers, inspect reproduction artifacts (code, data, outputs), rank attempts, and write structured feedback that helps teams improve model behavior and researcher-facing generation quality.
- You will evaluate document-type artifacts and provide evaluation ratings and written justifications.
- This work sits at the intersection of research expertise and careful, structured human evaluation.
The role (hours, type, language)
Position: ML Research Paper Reproduction Evaluator (contractor, part-time). Time commitment: 20+ hours/week. Remote and worldwide; fluency in English is required. OpenTrain AI is the contracting organization for this project.
- Employment types: Contractor, Part-time.
- Language: English required.
- Work format: Evaluate document artifacts and submit structured ratings and written feedback.
What you'll do
You will read original ML research papers, inspect AI-generated reproduction artifacts, evaluate reproducibility and methodological soundness, and produce ranked evaluations with clear justifications. Your feedback will guide improvements to AI agents that attempt to reproduce research.
- Read assigned ML research papers and understand contributions, methods, and results.
- Inspect reproduction artifacts including code, data, and model outputs.
- Evaluate technical correctness, completeness, and methodological soundness of reproductions.
- Compare and rank multiple reproduction attempts for the same paper.
- Provide clear, structured written justifications for rankings and evaluations.
Requirements
This is a specialist evaluator role with explicit academic and research requirements. Candidates must meet the core qualifications listed below and demonstrate strong research judgment and written communication.
- PhD in Machine Learning, AI, Computer Science, Statistics, or a closely related field (required).
- 2+ publications in top-tier ML/AI conferences such as NeurIPS, ICML, or ICLR (required).
- Strong research experience in areas like deep learning, generative AI, LLMs, representation learning, RL, optimization, computer vision, or NLP.
- Experience reading and critically evaluating research papers for technical and methodological correctness.
- Strong mathematical and statistical foundations.
- Excellent structured written communication and remote collaboration skills.
Helpful background and technical skills
Candidates with hands-on research or engineering experience will be most effective at this work. You should be comfortable inspecting code and experimental outputs and explaining issues clearly in writing.
- Experience as a research scientist, ML researcher, PhD researcher, or university faculty is helpful.
- Strong Python and familiarity with PyTorch or TensorFlow are valuable when inspecting reproduction code.
- Comfort judging experimental design, baselines, metrics, and statistical validity.
Who should apply
Apply if you are an experienced ML researcher who wants part-time, remote contract work that leverages your paper-reading and experimental-evaluation skills. This role is ideal for researchers who enjoy critical analysis and writing structured feedback that improves AI-generated research artifacts.
- Ideal for PhD researchers, research scientists, or faculty with ML publications.
- Good fit for people seeking 20+ hours/week of specialized remote contract work.
How the work is evaluated and delivered
You will receive assignments through OpenTrain, review documents and artifacts, submit evaluation ratings (EVALUATION_RATING) and written justifications, and collaborate remotely with reviewers or project managers. OpenTrain supports contributors in tracking assignments and building a portfolio of evaluations.
- Data type: DOCUMENT (papers, code, outputs).
- Labeling task: Evaluation rating plus structured written feedback.
- No pay rate is specified in this posting; compensation and contracting terms are handled through OpenTrain.