Skip to content
OpenTrain AIFor AI Companies

Machine Learning Engineer, Benchmarking & Evaluation

Join OpenTrain as a remote Machine Learning Engineer focused on benchmark-driven evaluation of real-world ML systems. This contractor role requires 3+ years of ML engineering experience, strong Python skills, and availability 20+ hrs/week with PST overlap.

OpenTrain AI

Coding & Software

Remote

10 countries

Eligibility

Intermediate

Experience

Jul 16, 2026

Posted

Open to applicants in

India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Brazil Mexico

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the leading platform for people building careers in AI training and data labeling. We help contributors discover projects, build a unified portfolio, and grow a durable freelance career in the human side of AI development.

We hire and contract directly for specialized AI training roles so you work on real problems that influence state-of-the-art systems while keeping the flexibility of remote, freelance work.

Why AI training and evaluation work matters

AI training (also called data labeling or human feedback work) is the human foundation behind modern machine learning. Evaluators and engineers prepare datasets, define metrics, and probe models so AI systems behave reliably in the real world.

This kind of work is highly flexible and accessible, and it gives you an opportunity to shape model behavior by working directly on benchmarking, validation, and failure-mode analysis.

The role

You will work hands-on with production-grade ML codebases to design and run benchmark-style evaluations. This is an engineering-heavy evaluation role: build and modify training, evaluation, and inference pipelines; prepare datasets and metrics; debug production-like systems; and analyze model behavior and edge cases.

Expect to write clean, reproducible Python code, partner with researchers and engineers, and deliver well-documented, repeatable evaluation workflows.

What you'll do

  • Work with real-world ML codebases on benchmark-style evaluation tasks.
  • Build, run, and modify model training, evaluation, and inference pipelines.
  • Prepare datasets, features, and metrics for ML benchmarking and validation.
  • Debug, refactor, and improve production-like ML systems for correctness and performance.
  • Evaluate model behavior, failure modes, and edge cases relevant to benchmark tasks.
  • Write clean, reproducible, and well-documented Python code for ML workflows.
  • Collaborate with researchers and engineers on challenging ML engineering tasks for AI system evaluation.

Requirements

You must meet the stated technical and availability requirements below.

  • 3+ years of experience as a Machine Learning Engineer or ML-focused Software Engineer.
  • Strong Python skills for machine learning and data workflows.
  • Hands-on experience building and modifying model training, evaluation, and inference pipelines.
  • Familiarity with ML fundamentals: supervised/unsupervised learning, evaluation metrics, and optimization.
  • Experience with ML frameworks such as PyTorch, TensorFlow, or JAX.
  • Ability to understand, navigate, and modify complex real-world ML codebases.
  • Strong problem-solving, debugging, and production-quality coding skills.
  • Excellent spoken and written English communication skills.

Helpful background

  • Experience participating in code reviews on ML engineering teams.
  • Comfort working in deployment-oriented workflows and production-like environments.

Logistics: schedule, contract, and location

This is a fully remote contractor assignment with a minimum commitment of 20 hours per week. The initial contract is 3 months and may be adjusted based on engagement.

You must be available at least 4 hours per day and have at least 4 hours of overlap with Pacific Standard Time (PST). English fluency is required.

  • Employment type: Contractor, Part-time.
  • Minimum time requirement: 20+ hours/week; at least 4 hours/day with 4 hours overlap with PST.
  • Contract length: 3 months (adjustable).
  • Eligible locations: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, Mexico.

How to apply

Apply through your OpenTrain account and submit a brief summary of relevant ML engineering experience, links to code or public projects if available, and your typical weekly availability in PST overlap hours.

We evaluate candidates on technical background, code experience with real ML systems, and clear communication. Shortlisted candidates may be asked to complete a technical exercise or code review task.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Machine Learning Benchmark Evaluator

Experienced ML engineers wanted for part-time, remote contract work evaluating production-grade model training, evaluation, and inference pipelines; requires 3+ years of ML engineering experience, strong Python, and PyTorch/TensorFlow/JAX familiarity. 20+ hours/week; apply through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +7 more
English
Part-time · Flexible
Expert level

Posted Jul 17, 2026

SWE Bench Data Engineer / Data Scientist — Benchmark Evaluation

Join OpenTrain AI as a contract, part-time SWE Bench contributor working 20+ hours/week on benchmark-driven data engineering and data-science evaluation. Use Python to build reproducible pipelines, process structured and unstructured data, and validate real-world workflows (English required).

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +7 more
English
Part-time · Flexible
Intermediate level

Posted Jul 16, 2026

AI Code Evaluation and Benchmarking Engineer

Evaluate and benchmark AI-generated code: review correctness, debug and verify solutions, and build evaluation datasets for frontier models. US-remote, contractor role — 20+ hrs/week (min 4 hrs/day), 1-month contract with 4-hour PST overlap required.

Coding & Software
Text
Remote · United States
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026