Skip to content
OpenTrain AIFor AI Companies

Data Engineering AI Evaluation Engineer

Build and validate Python data pipelines and benchmark tasks for advanced AI systems in a flexible, 3-month contractor role with 20+ hours per week.

OpenTrain AI

Coding & Software

Remote

10 countries

Eligibility

Entry

Experience

Jul 16, 2026

Posted

Open to applicants in

India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Brazil Mexico

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring and contracting for this role, connecting skilled contributors with hands-on projects that help shape how advanced AI systems are built and evaluated.

Creating an OpenTrain account is free, and applicants can apply in minutes while building a professional profile focused on AI training and data work.

About AI Evaluation Work

AI evaluation is the human side of improving artificial intelligence. Engineers and data specialists prepare datasets, create benchmarks, review outputs, and validate workflows so developers can measure whether AI systems perform accurately and reliably.

This work brings together software engineering, data science, and machine learning in a fast-growing field. Contributors work with production-like problems and help shape the behavior and capabilities of cutting-edge AI systems.

The Role

OpenTrain is recruiting a Data Engineering and Data Science AI Evaluation Engineer to design and validate data pipelines and evaluation tasks used to benchmark advanced AI systems. This is a hands-on contractor position involving production-like datasets, Python code, and real-world data workflows.

You will help create challenging, realistic tasks for AI systems and evaluate the quality, correctness, and reproducibility of the workflows and outputs used to measure model performance.

  • Contractor and part-time engagement
  • Three-month contract, adjustable based on engagement
  • At least 4 hours per day and 20 hours per week
  • Four hours of daily overlap with Pacific Standard Time
  • Available to candidates in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Türkiye, Brazil, and Mexico
  • English language communication required

What You'll Do

You will work across data engineering and data science workflows, using Python and complex real-world codebases to build reliable benchmarking and evaluation tasks. The role combines implementation, analysis, validation, documentation, and collaboration with researchers and engineers.

  • Work with structured and unstructured datasets for SWE Bench-style evaluation tasks
  • Design, build, and validate data pipelines for benchmarking and evaluation workflows
  • Perform data processing, analysis, feature preparation, and validation for data science use cases
  • Write, run, and modify Python code locally to process data and support experiments
  • Evaluate data quality, transformations, and outputs for correctness and reproducibility
  • Create clean, documented, and reusable workflows suitable for benchmarking
  • Participate in code reviews focused on quality and maintainability
  • Collaborate with researchers and engineers to design challenging, real-world AI tasks

Requirements

This role requires at least three years of overall experience as a Data Engineer, Data Scientist, or data-focused Software Engineer. You should be comfortable applying data engineering and data science methods in Python and working independently with complex technical problems.

  • At least 3 years of experience as a Data Engineer, Data Scientist, or data-focused Software Engineer
  • Strong proficiency in Python for data engineering and data science workflows
  • Demonstrable experience with data processing, analysis, and model-related workflows
  • Solid understanding of machine learning and data science fundamentals
  • Experience with structured and unstructured data
  • Ability to understand, navigate, and modify complex, real-world codebases
  • Experience writing readable, reusable, maintainable, and well-documented code
  • Strong problem-solving skills with algorithmic or data-intensive problems
  • Excellent spoken and written English communication skills

Helpful Background

Experience with benchmark-driven evaluation can help you contribute quickly, particularly when designing tasks that reflect realistic engineering and data science challenges.

  • Experience with SWE Bench or similar benchmark-driven evaluation projects
  • Background designing AI model evaluation tasks

Why Work in AI Training

AI training and data labeling work is a growing way to work in technology without needing to build an entire AI system alone. Specialists contribute directly by preparing examples, testing model behavior, evaluating results, and improving the datasets and workflows behind modern AI.

  • Remote work with a computer and internet connection
  • Flexible part-time opportunities that can fit around other commitments
  • Direct involvement with state-of-the-art AI systems
  • Opportunities to build experience in an expanding technical field

How to Apply

Create a free OpenTrain account, build your profile around your data engineering, data science, Python, and evaluation experience, and apply through OpenTrain. Be prepared to demonstrate your ability to work with complex codebases, production-like data, and reproducible evaluation workflows.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Software Engineering AI Evaluation Specialist

Evaluate AI-powered developer workflows through hands-on coding environments, GitHub, CI/CD, and technical assessment. This worldwide contractor role pays $50-$70 per hour for 20+ hours weekly.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $50–$70/hr

Posted Aug 19, 2026

Software Engineering AI Evaluation Expert

Use your software engineering judgment to evaluate AI-generated technical content, refine prompts, fact-check claims, and create high-quality engineering artifacts. Work remotely as a flexible contractor for 20+ hours per week.

Coding & Software
Document
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $100–$200/hr

Posted Aug 4, 2026

Frontend Engineering AI Evaluation Specialist

Use your frontend engineering expertise to evaluate code, architecture, and technical decisions that help improve AI systems. This remote contractor role offers $90-$140 per hour and requires 20+ hours each week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $90–$140/hr

Posted Sep 2, 2026