Join OpenTrain to evaluate AI-generated and human-created data science work: design grading criteria, score technical deliverables, and write defensible evaluations. Remote, contract role for experienced data scientists with 20+ hours/week availability and $100–$150/hr pay.
Generative AI & RLHF
100% Remote Hourly · $100–$150/hr
$100–$150/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Jul 29, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the leading platform for people building careers in AI training and data labeling. We help experienced contributors find specialized projects, consolidate opportunities, and build a unified portfolio that showcases their AI-training skills.
Creating an OpenTrain account is free. Our community connects skilled practitioners with evaluative work that directly shapes how AI systems perform in real-world tasks.
Platform for careers in AI training and data labeling
Free account and unified portfolio for contributors
Opportunities suitable for experienced data scientists
About AI training and evaluation work
AI training is the human side of building intelligent systems: people create, review, and score examples that models learn from. Evaluation work measures how well AI systems perform complex, technical tasks and guides model improvements.
Evaluators translate technical judgment into reproducible scoring, helping research teams compare model outputs, detect weaknesses, and prioritize fixes.
100% remote, flexible, and accessible work
Closely influences how models behave on real-world data science tasks
Often part-time contractor work that fits around other commitments
The role: Data Science AI Evaluation Expert
OpenTrain is recruiting experienced data scientists to evaluate AI-generated and human-created data science deliverables. You will design task-specific criteria, score outputs consistently, and provide written justifications that are evidence-based and reproducible.
This is a contract, part-time role that requires significant domain expertise and attention to detail.
Hours: 20+ hours/week
Pay: USD $100–$150 per hour (hourlyRate provided)
Work type: Contractor, Part-time; remote and open to applicants worldwide
Language: English; label type: EVALUATION_RATING; data type: TEXT
What you'll do day-to-day
You will translate technical expectations into clear grading rubrics and apply those rubrics to assess deliverables. Evaluations must be reproducible, defensible, and accompanied by concise written rationale.
You will also iterate on assessments when receiving structured feedback from senior reviewers and help refine evaluation criteria over time.
Design precise, task-specific grading criteria for EDA, modeling, ML pipelines, experiments, feature engineering, and technical reports
Score AI-generated or human-created data science work against established rubrics
Provide detailed written justifications for each evaluation and score
Apply evidence-based judgment to ensure reproducibility and consistency
Incorporate structured feedback from senior reviewers and revise assessments as needed
Requirements
You must have professional data science experience and a track record at a leading technology, research, or quantitative firm. Strong technical skills and excellent technical writing are essential.
All required qualifications below come from the role description; applicants who meet them will be strongly considered.
Minimum 1+ years of professional data science experience
Experience at a leading technology, research, or quantitative firm (e.g., FAANG, AI labs, top-tier quant funds)
Strong command of Python and SQL
Proficiency in statistical modeling, machine learning, experimentation, and causal inference
Exceptional written communication for conveying technical findings
Detail-oriented and consistent approach to evaluating complex work
How it works and how to apply
Apply through your OpenTrain account — creating one is free. We will review applications for relevant experience and technical fit, then onboard approved evaluators with task instructions and rubric examples.
Expect to perform evaluations independently, submit scores and justifications, and respond to reviewer feedback as part of an iterative quality process.
Create a free OpenTrain account and submit your profile and experience
Applications are reviewed for domain experience and technical fit
Onboarding provides rubric templates, scoring examples, and reviewer feedback loops
Design and author multi-step scientific evaluation tasks for frontier AI models in a full-time remote US contractor role paying $60–$90/hr. Expect ~35 hours/week building Python reference solutions, defining rigorous criteria, and reviewing model attempts.
Use your computational biology expertise to evaluate and annotate AI-generated outputs across genomics, structural biology, and systems biology. Remote contractor role, 20+ hrs/week, $40–$60/hr — help shape safer, more accurate scientific AI.
Join OpenTrain as a contractor designing and validating data pipelines and benchmark evaluation tasks for AI systems. Work part-time (20+ hrs/wk) with Python on production-like datasets; applicants need 3+ years in data engineering, data science, or data-focused software engineering.