AI Data Scientist, Model Evaluation & Reference Solutions
Join OpenTrain as a remote AI Data Scientist to review AI-generated analysis and code, write step-by-step reference solutions, and rate model outputs. Part-time contractor work (20+ hrs/week) with hourly pay up to $100 and hiring across many countries.
Generative AI & RLHF
Remote Hourly · $100/hr
$100/hr
Compensation
17 countries
Eligibility
Entry
Experience
Jul 8, 2026
Posted
Open to applicants in
Bangladesh Bhutan Brazil Cambodia Germany India Indonesia Malaysia Nepal Pakistan Singapore Sri Lanka Thailand Philippines United States Timor-Leste Vietnam
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people who build careers in AI training and data labeling. We help contributors discover projects, consolidate opportunities, and build a unified portfolio they control. Creating an OpenTrain account is free.
We connect skilled contractors to ongoing model-evaluation and annotation work so you can grow from single projects into a durable freelance career in the fast-growing AI training industry.
About AI training work
AI training (also called data labeling, annotation, or human feedback work) is the human side of how modern AI learns. Experts evaluate model outputs, rate responses, produce reference answers, and identify errors so models improve over time.
This role focuses on training and evaluating models that produce analytical reasoning, code, and quantitative conclusions. Your work will directly shape how data-driven AI systems reason and explain their results.
Role overview
You will review AI-generated analytical reasoning, code, and model outputs, produce clear reference solutions, and judge which AI responses are most correct and well reasoned. This is a remote, contractor position built around model evaluation and training content focused on data science and machine learning tasks.
Labeling work will include RLHF-style evaluation, qualitative ratings of reasoning, and creating training examples for text generation and evaluation tasks.
Employment type: Contractor, part-time.
Time commitment: 20+ hours per week.
Labeling tasks include: RLHF, evaluation rating, and text generation.
What you'll do
Review AI-generated analytical reasoning, code, and model outputs for accuracy, clarity, and methodological soundness.
Write step-by-step reference solutions and clear explanations for complex quantitative and data problems.
Fact-check quantitative claims and identify statistical or experimental-design errors.
Rate and compare multiple AI responses on correctness, reasoning quality, and adherence to prompts.
Create detailed prompts and exemplar responses to improve model learning across varied data topics.
Test models for inaccuracies, bias, and inconsistent behavior; validate model reliability on representative cases.
Requirements
Bachelor’s degree or higher in Data Science, Computer Science, Statistics, Mathematics, or a closely related quantitative field.
5+ years of professional experience as a Data Scientist or in a closely related analytical role.
Strong Python skills for data analysis and machine learning, including pandas, NumPy, and scikit-learn.
Solid background in statistics, experimental design, and applied probability.
Hands-on experience building, evaluating, and deploying machine learning models.
Advanced SQL skills and comfort working with large, complex datasets.
Minimum C1 English proficiency with the ability to write clear quantitative explanations.
Experience with data visualization tools or dashboards and presenting findings to stakeholders.
Prior experience reviewing AI-generated analytical content, annotation, or model-evaluation work is a strong plus.
Highly detail-oriented and systematic, with a focus on quality, reproducibility, and careful evaluation of reasoning steps.
Work setup, pay, and locations
This is a remote contractor engagement with hourly pay. Typical assignments are part-time and flexible but expect a baseline of 20+ hours per week.
Compensation: hourly pay up to $100. Country-based variants include up to $50, up to $70, and up to $100 per hour depending on location.
Remote hiring across: Vietnam, Timor Leste, Thailand, Singapore, Pakistan, The Philippines, Nepal, Malaysia, Sri Lanka, Cambodia, Indonesia, Bhutan, Bangladesh, India, Brazil, Germany, and the United States.
Access to future projects and related model-evaluation work as opportunities arise.
Who should apply and next steps
Apply if you are an experienced data practitioner who enjoys explaining quantitative work, assessing model reasoning, and producing reproducible reference solutions. Strong Python, statistics, SQL, and hands-on ML experience are essential.
To apply, create a free OpenTrain account and submit your application for this role. Qualified applicants will be contacted with next steps and onboarding information.
Review AI-generated SQL queries and database designs, write step-by-step reference solutions, and rate model responses — remote contractor role at $50/hour, 20+ hrs/week, open to qualified applicants in approved countries.
Join OpenTrain as a contractor designing and validating data pipelines and benchmark evaluation tasks for AI systems. Work part-time (20+ hrs/wk) with Python on production-like datasets; applicants need 3+ years in data engineering, data science, or data-focused software engineering.
Join OpenTrain as an AI Model Evaluation Developer to write and maintain code, run model benchmarks, rank responses, and build datasets for fine-tuning and RLHF. This remote, part-time contractor role requires strong Python and JavaScript/TypeScript skills and 20+ hours/week.