Join OpenTrain as an R Model Evaluation Engineer to review AI-generated analyses, write expert R solutions, and improve model quality. Remote contract work at $55/hr, 20+ hours/week, for experienced R users who can judge statistical correctness and explain decisions clearly.
Generative AI & RLHF
Remote Hourly · $55/hr
$55/hr
Compensation
17 countries
Eligibility
Intermediate
Experience
Jul 9, 2026
Posted
Open to applicants in
Bangladesh Bhutan Brazil Cambodia Germany India Indonesia Malaysia Nepal Pakistan Singapore Sri Lanka Thailand Philippines United States Timor-Leste Vietnam
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people who build careers in AI training and data labeling. We help contributors find projects, consolidate their work history, and grow a durable freelance career teaching AI through careful, accountable annotation and evaluation.
As the hiring organization for this role, OpenTrain offers flexible, remote contract work that connects your R and statistics expertise with meaningful model-improvement tasks across cutting-edge AI systems.
About AI training work
AI training (also called data labeling or human feedback work) is the human side of building intelligent systems: people annotate, evaluate, and improve model outputs so AI behaves more reliably and usefully. These opportunities are often remote, flexible, and accessible to contributors with strong domain skills rather than prior platform experience.
This role focuses on evaluating analytical outputs and model reasoning in R—one of the most important practical skills shaping how data-driven AI responds to statistical and data-wrangling tasks.
The role
You will evaluate AI-generated responses, craft gold-standard R solutions, and judge correctness, clarity, and reasoning quality for data-analytic prompts. This is a remote, hourly contract position (part-time) requiring deep, hands-on R experience and applied statistics knowledge.
Hourly rate: $55 USD per hour.
Time commitment: 20+ hours per week.
Employment type: Contractor, Part-time.
Work language: English (en).
Eligible countries: BD, BT, BR, KH, DE, IN, ID, MY, NP, PK, SG, LK, TH, PH, US, TL, VN.
What you'll do
Review AI-generated responses for accuracy, clarity, adherence to prompts, and step-by-step reasoning quality.
Identify errors in statistical methods, modeling choices, and data-wrangling workflows.
Fact-check analytical results and compare multiple model responses for correctness.
Write expert-level explanations and model solutions that demonstrate correct use of R.
Create detailed prompts and responses to support supervised fine-tuning and RLHF training.
Test model outputs for inaccuracies or biases and assess reliability across use cases.
Requirements
You must meet the core technical and communication requirements below; we will verify relevant experience during selection.
Minimum 2 years of hands-on experience using R for data analysis, statistics, or data science work.
Strong proficiency with R programming: data wrangling, functional programming patterns, and reusable functions or packages.
Solid grounding in applied statistics, including regression, inference, and model validation.
Experience building end-to-end analyses in R: data cleaning, exploratory analysis, modeling, and visualization.
Familiarity with tidyverse, data.table, and ggplot2.
Experience reviewing AI-generated or model-produced analyses for correctness and reasoning quality.
Excellent written English and the ability to explain analytical decisions clearly.
Bachelor’s degree in Statistics, Mathematics, Computer Science, or a closely related quantitative field.
Who should apply
This project is built for intermediate-level R practitioners who enjoy detailed code review and explaining statistical thinking in writing. If you routinely produce reproducible R analyses and can evaluate alternative model responses for correctness, you'll be a strong fit.
You should be comfortable working independently, managing part-time hours across a flexible schedule, and contributing precise examples and feedback that directly improve model behavior.
How it works — logistics and next steps
OpenTrain hires contributors as contractors for remote, hourly work. This role uses text-based labeling and includes RLHF-style evaluation, rating, and prompt/response writing tasks. Compensation is $55/hour and the schedule is 20+ hours/week.
To apply, create a free OpenTrain account, complete your profile, and submit your application. If selected, you'll receive onboarding materials, example tasks, and quality guidelines to get started. OpenTrain centralizes your work history so you can build a lasting portfolio in AI training.
Labeling types for this role: RLHF, evaluation rating, and prompt/response writing (SFT).
Data type: text; primary subject matter: R data analysis and applied statistics.
Join OpenTrain AI to evaluate LLM outputs in finance, design rubrics, and help shape model training and benchmarks. Part-time contractor role (<20 hrs/week), remote worldwide, paying $100/hr for finance professionals with 2+ years' experience.
Join OpenTrain as an AI Model Evaluation Developer to write and maintain code, run model benchmarks, rank responses, and build datasets for fine-tuning and RLHF. This remote, part-time contractor role requires strong Python and JavaScript/TypeScript skills and 20+ hours/week.
Join OpenTrain AI as a Machine Learning Engineer to design and run evaluation benchmarks for frontier LLMs. Contract, part-time role (20+ hrs/week) paying $60–$90/hr for US-based practitioners with hands-on ML training and research experience.