Join OpenTrain as a remote Machine Learning Engineer focused on benchmark-driven evaluation of real-world ML systems. This contractor role requires 3+ years of ML engineering experience, strong Python skills, and availability 20+ hrs/week with PST overlap.
Coding & Software
Remote
10 countries
Eligibility
Intermediate
Experience
Jul 16, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Brazil Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the leading platform for people building careers in AI training and data labeling. We help contributors discover projects, build a unified portfolio, and grow a durable freelance career in the human side of AI development.
We hire and contract directly for specialized AI training roles so you work on real problems that influence state-of-the-art systems while keeping the flexibility of remote, freelance work.
Why AI training and evaluation work matters
AI training (also called data labeling or human feedback work) is the human foundation behind modern machine learning. Evaluators and engineers prepare datasets, define metrics, and probe models so AI systems behave reliably in the real world.
This kind of work is highly flexible and accessible, and it gives you an opportunity to shape model behavior by working directly on benchmarking, validation, and failure-mode analysis.
The role
You will work hands-on with production-grade ML codebases to design and run benchmark-style evaluations. This is an engineering-heavy evaluation role: build and modify training, evaluation, and inference pipelines; prepare datasets and metrics; debug production-like systems; and analyze model behavior and edge cases.
Expect to write clean, reproducible Python code, partner with researchers and engineers, and deliver well-documented, repeatable evaluation workflows.
What you'll do
Work with real-world ML codebases on benchmark-style evaluation tasks.
Build, run, and modify model training, evaluation, and inference pipelines.
Prepare datasets, features, and metrics for ML benchmarking and validation.
Debug, refactor, and improve production-like ML systems for correctness and performance.
Evaluate model behavior, failure modes, and edge cases relevant to benchmark tasks.
Write clean, reproducible, and well-documented Python code for ML workflows.
Collaborate with researchers and engineers on challenging ML engineering tasks for AI system evaluation.
Requirements
You must meet the stated technical and availability requirements below.
3+ years of experience as a Machine Learning Engineer or ML-focused Software Engineer.
Strong Python skills for machine learning and data workflows.
Hands-on experience building and modifying model training, evaluation, and inference pipelines.
Familiarity with ML fundamentals: supervised/unsupervised learning, evaluation metrics, and optimization.
Experience with ML frameworks such as PyTorch, TensorFlow, or JAX.
Ability to understand, navigate, and modify complex real-world ML codebases.
Strong problem-solving, debugging, and production-quality coding skills.
Excellent spoken and written English communication skills.
Helpful background
Experience participating in code reviews on ML engineering teams.
Comfort working in deployment-oriented workflows and production-like environments.
Logistics: schedule, contract, and location
This is a fully remote contractor assignment with a minimum commitment of 20 hours per week. The initial contract is 3 months and may be adjusted based on engagement.
You must be available at least 4 hours per day and have at least 4 hours of overlap with Pacific Standard Time (PST). English fluency is required.
Employment type: Contractor, Part-time.
Minimum time requirement: 20+ hours/week; at least 4 hours/day with 4 hours overlap with PST.
Apply through your OpenTrain account and submit a brief summary of relevant ML engineering experience, links to code or public projects if available, and your typical weekly availability in PST overlap hours.
We evaluate candidates on technical background, code experience with real ML systems, and clear communication. Shortlisted candidates may be asked to complete a technical exercise or code review task.
Experienced ML engineers wanted for part-time, remote contract work evaluating production-grade model training, evaluation, and inference pipelines; requires 3+ years of ML engineering experience, strong Python, and PyTorch/TensorFlow/JAX familiarity. 20+ hours/week; apply through OpenTrain.
Join OpenTrain AI as a contract, part-time SWE Bench contributor working 20+ hours/week on benchmark-driven data engineering and data-science evaluation. Use Python to build reproducible pipelines, process structured and unstructured data, and validate real-world workflows (English required).