Join OpenTrain as a Machine Learning Engineer evaluating real-world ML systems through benchmark-driven coding tasks. Work remotely on training, inference, debugging, and model evaluation for at least 20 hours per week.
Coding & Software
Remote
10 countries
Eligibility
Intermediate
Experience
Jul 16, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Brazil Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and apply in minutes.
Creating an OpenTrain account is free, and this opportunity offers a way to contribute directly to the development and evaluation of advanced AI systems.
Fully remote contractor assignment
Part-time engagement of at least 20 hours per week
Open to candidates in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, and Mexico
About AI Training and Model Evaluation
AI training is the human side of building artificial intelligence. Engineers, reviewers, and other specialists prepare data, test model behavior, evaluate outputs, and identify failures that help make AI systems more capable and reliable.
In this role, your machine learning engineering expertise will support benchmark-driven evaluation work involving real-world codebases, model pipelines, datasets, metrics, and production-like systems.
Work on cutting-edge AI development and evaluation
Use engineering judgment to assess correctness, performance, and edge cases
Contribute remotely with a flexible part-time schedule
The Role
OpenTrain is recruiting an experienced Machine Learning Engineer for hands-on benchmark-style evaluation tasks. You will build and modify training, evaluation, and inference workflows; prepare benchmarking materials; and investigate how machine learning systems perform across challenging scenarios.
The work requires clean, reproducible Python code, strong debugging ability, and the confidence to navigate complex ML codebases. You will also collaborate with researchers and engineers on difficult engineering tasks connected to AI system evaluation.
Role focus: machine learning engineering and benchmark-driven evaluation
Experience level: intermediate
Primary language: English
Contract length: 3 months, adjustable based on engagement
What You'll Do
You will work across the development and evaluation lifecycle of machine learning systems. Assignments may involve implementing changes, running experiments, preparing validation materials, and diagnosing behavior in production-like environments.
Work with real-world ML codebases on benchmark-style evaluation tasks
Build, run, and modify model training, evaluation, and inference pipelines
Prepare datasets, features, and metrics for ML benchmarking and validation
Debug, refactor, and improve production-like ML systems for correctness and performance
Evaluate model behavior, failure modes, and edge cases relevant to benchmark tasks
Write clean, reproducible, and well-documented Python code for ML workflows
Collaborate with researchers and engineers on challenging ML engineering tasks for AI system evaluation
Required Skills and Experience
This opportunity is designed for an experienced machine learning or ML-focused software engineer who can independently understand complex systems and produce reliable engineering work. Strong English communication is required for written and spoken collaboration.
At least 3 years of experience as a Machine Learning Engineer or ML-focused Software Engineer
Strong Python skills for machine learning and data workflows
Hands-on experience with model training, evaluation, and inference pipelines
Familiarity with supervised and unsupervised learning, evaluation metrics, and optimization
Experience with ML frameworks such as PyTorch, TensorFlow, or JAX
Ability to understand, navigate, and modify complex real-world ML codebases
Strong problem-solving, debugging, and production-quality coding skills
Excellent spoken and written English communication skills
Helpful Background
Experience working within disciplined engineering environments will help you contribute effectively to benchmark and evaluation assignments.
Experience with code reviews in ML engineering teams
Comfort working in deployment-oriented workflows and production-like environments
Clear English communication and production-quality coding habits
Work Arrangement and Application
This is a fully remote, part-time contractor assignment requiring a minimum of 20 hours per week. You must be available for at least 4 hours per day, including 4 hours of overlap with Pacific Standard Time.
Apply through OpenTrain by creating a free account and submitting your profile. OpenTrain brings together opportunities in AI training and data labeling so you can build experience in a fast-growing technical field.
Minimum commitment: 20+ hours per week
Daily availability: at least 4 hours
Required schedule overlap: 4 hours with PST
Contractor and part-time arrangement
Eligible locations: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, and Mexico
Evaluate AI agents on realistic software engineering tasks and build reliable MCP-based reinforcement learning environments. This flexible remote contractor role offers an ideal rate range of $60 to $120 per hour.
Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.
Build and improve machine learning models remotely as a contractor using Python, MongoDB, major ML frameworks, and large datasets. Earn $80-$140 per hour while contributing to practical AI training workflows.