Skip to content
OpenTrain AIFor AI Companies

Machine Learning Engineer MLE Bench Evaluation

Join OpenTrain as a Machine Learning Engineer evaluating real-world ML systems through benchmark-driven coding tasks. Work remotely on training, inference, debugging, and model evaluation for at least 20 hours per week.

OpenTrain AI

Coding & Software

Remote

10 countries

Eligibility

Intermediate

Experience

Jul 16, 2026

Posted

Open to applicants in

India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Brazil Mexico

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and apply in minutes.

Creating an OpenTrain account is free, and this opportunity offers a way to contribute directly to the development and evaluation of advanced AI systems.

  • Fully remote contractor assignment
  • Part-time engagement of at least 20 hours per week
  • Open to candidates in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, and Mexico

About AI Training and Model Evaluation

AI training is the human side of building artificial intelligence. Engineers, reviewers, and other specialists prepare data, test model behavior, evaluate outputs, and identify failures that help make AI systems more capable and reliable.

In this role, your machine learning engineering expertise will support benchmark-driven evaluation work involving real-world codebases, model pipelines, datasets, metrics, and production-like systems.

  • Work on cutting-edge AI development and evaluation
  • Use engineering judgment to assess correctness, performance, and edge cases
  • Contribute remotely with a flexible part-time schedule

The Role

OpenTrain is recruiting an experienced Machine Learning Engineer for hands-on benchmark-style evaluation tasks. You will build and modify training, evaluation, and inference workflows; prepare benchmarking materials; and investigate how machine learning systems perform across challenging scenarios.

The work requires clean, reproducible Python code, strong debugging ability, and the confidence to navigate complex ML codebases. You will also collaborate with researchers and engineers on difficult engineering tasks connected to AI system evaluation.

  • Role focus: machine learning engineering and benchmark-driven evaluation
  • Experience level: intermediate
  • Primary language: English
  • Contract length: 3 months, adjustable based on engagement

What You'll Do

You will work across the development and evaluation lifecycle of machine learning systems. Assignments may involve implementing changes, running experiments, preparing validation materials, and diagnosing behavior in production-like environments.

  • Work with real-world ML codebases on benchmark-style evaluation tasks
  • Build, run, and modify model training, evaluation, and inference pipelines
  • Prepare datasets, features, and metrics for ML benchmarking and validation
  • Debug, refactor, and improve production-like ML systems for correctness and performance
  • Evaluate model behavior, failure modes, and edge cases relevant to benchmark tasks
  • Write clean, reproducible, and well-documented Python code for ML workflows
  • Collaborate with researchers and engineers on challenging ML engineering tasks for AI system evaluation

Required Skills and Experience

This opportunity is designed for an experienced machine learning or ML-focused software engineer who can independently understand complex systems and produce reliable engineering work. Strong English communication is required for written and spoken collaboration.

  • At least 3 years of experience as a Machine Learning Engineer or ML-focused Software Engineer
  • Strong Python skills for machine learning and data workflows
  • Hands-on experience with model training, evaluation, and inference pipelines
  • Familiarity with supervised and unsupervised learning, evaluation metrics, and optimization
  • Experience with ML frameworks such as PyTorch, TensorFlow, or JAX
  • Ability to understand, navigate, and modify complex real-world ML codebases
  • Strong problem-solving, debugging, and production-quality coding skills
  • Excellent spoken and written English communication skills

Helpful Background

Experience working within disciplined engineering environments will help you contribute effectively to benchmark and evaluation assignments.

  • Experience with code reviews in ML engineering teams
  • Comfort working in deployment-oriented workflows and production-like environments
  • Clear English communication and production-quality coding habits

Work Arrangement and Application

This is a fully remote, part-time contractor assignment requiring a minimum of 20 hours per week. You must be available for at least 4 hours per day, including 4 hours of overlap with Pacific Standard Time.

Apply through OpenTrain by creating a free account and submitting your profile. OpenTrain brings together opportunities in AI training and data labeling so you can build experience in a fast-growing technical field.

  • Minimum commitment: 20+ hours per week
  • Daily availability: at least 4 hours
  • Required schedule overlap: 4 hours with PST
  • Contractor and part-time arrangement
  • Eligible locations: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, and Mexico

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

MCP AI Software Evaluation Engineer

Evaluate AI agents on realistic software engineering tasks and build reliable MCP-based reinforcement learning environments. This flexible remote contractor role offers an ideal rate range of $60 to $120 per hour.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $60–$120/hr

Posted Aug 25, 2026

C++ LLM Evaluation Software Engineer

Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026

Machine Learning Model Development Engineer

Build and improve machine learning models remotely as a contractor using Python, MongoDB, major ML frameworks, and large datasets. Earn $80-$140 per hour while contributing to practical AI training workflows.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $80–$140/hr

Posted Jul 3, 2026