Remote US-based contractor role reviewing AI-generated Python ML code and workflows — $120/hr, 20+ hours/week. Evaluate correctness, spot data leakage, give clear technical feedback, and help onboard and calibrate contributor teams.
Coding & Software
Remote Hourly · $120/hr
$120/hr
Compensation
1 country
Eligibility
Intermediate
Experience
Jul 9, 2026
Posted
Open to applicants in
United States
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people building careers in AI training and data labeling. We hire and contract contributors directly to do the human work that helps modern AI learn: annotating, evaluating, and reviewing model outputs across many technical domains.
Creating an OpenTrain account is free. Contributors use it to find projects, build a unified AI training portfolio, and grow into durable freelance careers working remotely.
About AI training work
AI training (data labeling, annotation, and human evaluation) is the human side of building modern models. Contributors review outputs, rate quality, correct errors, and provide the nuanced judgment machines still need.
This role focuses on Python machine-learning work — reviewing code, pipelines, notebooks, metrics, and explanations to ensure reproducibility and sound methodology.
100% remote and flexible work that fits alongside other commitments.
Accessible roles range from entry-level annotation to specialist technical review; this listing is a technical, intermediate-level QA role.
The role
OpenTrain is recruiting a Python ML Quality Assurance Lead to evaluate AI-generated Python code, ML workflows, and trainer QA output. You will monitor quality, provide written technical feedback, and help contributors uphold rigorous standards.
This is a US-only, part-time contractor role (20+ hours/week) at $120/hour. Work is remote and focuses on rubric-based review, onboarding support, and process improvement.
Employment type: Contractor, Part-time
Hours: 20+ hours/week
Rate: $120 USD per hour
Location: United States (remote)
Language: English required
What you'll do
You will review AI-generated Python ML artifacts and help keep reviewer output consistent, reproducible, and methodologically sound. Expect a mix of hands-on review, feedback writing, and contributor-facing documentation work.
Review AI-generated Python code, ML pipelines, preprocessing, training workflows, evaluation logic, and explanations for correctness and reproducibility.
Spot-check for data leakage, statistical validity, debugging accuracy, formatting, readability, and maintainability.
Evaluate whether metrics, model selection, cross-validation, and feature-engineering choices are appropriate.
Provide precise written feedback to contributors and escalate recurring or critical quality issues.
Update guidance, trackers, FAQs, examples, honeypots, and onboarding materials to improve reviewer calibration.
Support onboarding and training calls for Python ML contributors and help maintain scalable QA processes for remote teams.
Requirements
You must have strong Python fundamentals and practical machine-learning judgment. This is an intermediate-level technical reviewer role; leadership or team coordination experience is a plus.
Strong Python knowledge: data structures, functions, classes, iterators, comprehensions, exception handling, virtualenv/package management, testing, and debugging.
Strong ML knowledge: supervised/unsupervised methods, feature engineering, train/test splits, cross-validation, model selection, bias/variance, regularization, metrics, and reproducibility.
Able to evaluate ML content against detailed rubrics and identify flawed methodology, wrong metrics, hallucinated APIs, misleading conclusions, and incomplete explanations.
Practical experience spotting data leakage, invalid model assumptions, and flawed evaluation pipelines.
Clear written English communication for concise technical feedback and team coordination.
Preferred: prior experience with AI training, LLM evaluation, code QA, ML QA, rubric-based review, or leading remote technical contributors.
How it works & how to apply
OpenTrain is the hiring and contracting organization for this role. To apply, create a free OpenTrain account, complete your profile, and submit your application through the platform.
When you apply, highlight relevant Python and ML experience and samples (not required but helpful). Successful applicants will be invited to calibration exercises and onboarding to align on rubrics and review standards.
Prepare a short summary of relevant Python/ML work and availability (20+ hrs/week).
You will participate in calibration tasks and onboarding sessions to ensure alignment with project rubrics.
Review and rate AI-generated Python code for correctness, readability, security, and rubric adherence in remote US-based contract work; flexible 20+ hours/week at up to $70/hr. Join OpenTrain to help shape Python training quality and mentor contributors.
OpenTrain AI seeks experienced Python QA engineers to review code and ensure SLA-level quality for LLM data training projects. This remote, contractor role requires 5+ years of Python, a B2+ English level, and a reliable 30–40 hours/week commitment for 6+ months at $13/hr.
Join OpenTrain as an R Quality Assurance Lead to review R code, statistical analyses, visualizations, and reproducible workflows for correctness and clarity; contract, remote (US only), up to $65/hr, ~20+ hours/week. Help shape R-specific QA standards, documentation, and trainer support for high-qua