Skip to content
OpenTrain AIFor AI Companies

Machine Learning Engineering Evaluator

Apply now

Machine Learning Engineering Evaluator

Evaluate demanding machine learning engineering tasks involving models, training systems, inference, Python, and performance optimization. Work remotely worldwide for approximately 15 hours per week at $100-$150 per hour.

OpenTrain AI

Coding & Software

100% Remote Hourly · $100–$150/hr

$100–$150/hr

Compensation

Worldwide

Eligibility

Entry

Experience

Sep 11, 2026

Posted

Open worldwide

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps people discover projects, build a professional profile, and apply in minutes as they develop experience in this fast-growing field.

Creating an OpenTrain account is free. Your profile can help you present credible AI training experience and find opportunities that match your technical background.

  • Remote AI training and data-labeling opportunities
  • A profile for building a lasting AI training portfolio
  • Flexible work that can fit around other commitments

About AI Training Work

AI training is the human side of building artificial intelligence. Technical contributors create, test, review, and improve the examples, code, and evaluation systems that help modern AI models become more capable and reliable.

In this role, your engineering judgment will support model development and evaluation by checking correctness, reproducibility, efficiency, performance, and technical reasoning.

  • Work directly with cutting-edge machine learning systems
  • Evaluate implementations against objective technical requirements
  • Help improve how AI systems are developed and assessed

The Role

OpenTrain is recruiting a Machine Learning Engineering Evaluator to create, solve, review, and validate demanding machine learning engineering tasks for an AI training project. The work spans model development, training and inference systems, numerical computing, performance optimization, Python workflows, and technical evaluation.

This is a global, fully remote contractor role requiring approximately 15 hours per week. Scheduling is flexible, including the option to work weekends. Compensation is listed at $100-$150 per hour.

  • Work worldwide as a fully remote contractor
  • Work approximately 15 hours per week
  • Choose a flexible schedule, including weekends
  • Earn a listed rate of $100-$150 per hour
  • Work in English

What You'll Do

You will develop and validate machine learning models, training pipelines, inference systems, and supporting infrastructure. You will also assess AI-generated code and technical solutions, documenting technical decisions, trade-offs, limitations, and failure modes.

The work requires objective testing and careful diagnosis across model behavior, numerical operations, system performance, and reproducibility.

  • Implement model components, data pipelines, evaluation systems, and numerical methods
  • Build reproducible workflows with Python and command-line tools
  • Work with tensor operations, automatic differentiation, model architectures, tokenization, batching, and generation
  • Optimize latency, throughput, memory usage, and hardware utilization
  • Diagnose numerical instability, tensor errors, memory bottlenecks, distributed-system failures, and performance regressions
  • Review AI-generated code and technical solutions
  • Design objective tests, benchmarks, and verification criteria
  • Evaluate correctness, reproducibility, efficiency, and performance

Required Qualifications

You should have a master's degree or PhD in computer science, machine learning, artificial intelligence, applied mathematics, statistics, engineering, or a closely related quantitative discipline, along with strong professional or research experience in machine learning.

You must be able to debug machine learning systems beyond surface-level API usage and clearly explain implementation choices, performance trade-offs, and failure modes. Equivalent tools and substantial open-source or academic experience may also qualify.

  • Master's degree or PhD in a relevant quantitative discipline
  • Strong professional or research experience in machine learning
  • Practical proficiency with Python
  • Experience building reproducible technical workflows
  • Meaningful experience with at least two relevant ML frameworks, libraries, or inference tools
  • Strong understanding of model training, evaluation, numerical computation, or inference
  • Ability to debug ML systems beyond surface-level API usage
  • Ability to explain implementation decisions, performance trade-offs, and failure modes clearly

Relevant Technical Tools

Experience may include the tools below, along with equivalent technologies. The role is focused on practical machine learning engineering judgment rather than surface-level familiarity with APIs.

  • PyTorch
  • JAX
  • NumPy
  • SciPy
  • SGLang
  • vLLM
  • llama.cpp
  • Hugging Face Transformers
  • Hugging Face Tokenizers
  • Equivalent tools or substantial open-source and academic experience

Why Build an AI Training Career With OpenTrain

AI training and data-labeling work is a rapidly growing way to work in technology. Contributors with specialized expertise can help shape how state-of-the-art AI systems behave while developing a portfolio of practical, high-value experience.

OpenTrain gives you one place to build your profile, discover projects across the industry, and apply for work that matches your skills. This role offers an opportunity to turn machine learning engineering expertise into flexible, remote AI training work.

  • Work from anywhere with a computer and internet connection
  • Choose flexible part-time work around your schedule
  • Build a credible portfolio of AI training experience
  • Apply your machine learning engineering expertise to advanced AI systems

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Machine Learning Engineer MLE Bench Evaluation

Join OpenTrain as a Machine Learning Engineer evaluating real-world ML systems through benchmark-driven coding tasks. Work remotely on training, inference, debugging, and model evaluation for at least 20 hours per week.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +7 more
English
Part-time · Flexible
Intermediate level

Posted Jul 16, 2026

Data Engineering AI Evaluation Engineer

Build and validate Python data pipelines and benchmark tasks for advanced AI systems in a flexible, 3-month contractor role with 20+ hours per week.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +7 more
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026

AI Evaluation Engineer, Engineering Simulation

Create and validate challenging engineering simulation benchmarks that train and evaluate AI agents. This remote contractor role combines advanced engineering design, Python, open-source simulation, and model failure analysis.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Sep 11, 2026