Skip to content
OpenTrain AIFor AI Companies

SWE Bench Data Engineer / Data Scientist — Benchmark Evaluation

Join OpenTrain AI as a contract, part-time SWE Bench contributor working 20+ hours/week on benchmark-driven data engineering and data-science evaluation. Use Python to build reproducible pipelines, process structured and unstructured data, and validate real-world workflows (English required).

OpenTrain AI

Coding & Software

Remote

10 countries

Eligibility

Intermediate

Experience

Jul 16, 2026

Posted

Open to applicants in

India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Brazil Mexico

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the centralized, open platform where people build careers in AI training and data labeling. Contributors use OpenTrain to find projects, consolidate work history, and develop a reusable portfolio that demonstrates their skills across AI evaluation and annotation work.

OpenTrain AI is the hiring and contracting organization for this role. Creating an OpenTrain account is free and lets you apply, manage opportunities, and build a record of benchmark and annotation work that employers respect.

  • A focused career platform for people who train and evaluate AI systems
  • Work is remote, flexible, and directly shapes how modern AI behaves
  • Free to join and build a unified portfolio across projects

About AI Training Work

AI training (data labeling, annotation, and human evaluation) is the human work behind modern AI systems. Contributors annotate, validate, and review data and model outputs so researchers and engineers can measure and improve model behavior.

This role sits at the intersection of data engineering and data science: you will support benchmark-driven evaluations that mimic production workflows to ensure results are correct, reproducible, and useful for model development.

  • 100% remote, often flexible hours and part-time-friendly
  • Work ranges from beginner-friendly annotation to advanced engineering and evaluation
  • Contributors directly affect how state-of-the-art AI systems are assessed

The Role

We are hiring experienced SWE Bench Data Engineers / Data Scientists to perform benchmark-driven evaluation work focused on real-world data engineering and data science workflows. The work centers on structured and unstructured datasets, production-like data pipelines, feature preparation, data validation, and Python-based experimentation in complex codebases.

This is a contract, part-time role expecting 20+ hours per week. Candidates must be fluent in spoken and written English and be located in one of the approved countries: IN, PK, NG, KE, EG, GH, BD, TR, BR, MX. Experience level: Intermediate (3+ years).

  • Employment types: Contractor, Part-time
  • Time requirement: 20+ hours/week
  • Languages: English (spoken and written)
  • Eligible countries: IN, PK, NG, KE, EG, GH, BD, TR, BR, MX

What You'll Do

  • Build, validate, and maintain data pipelines for benchmarking and evaluation workflows
  • Process, analyze, and prepare structured and unstructured data for data-science use cases
  • Write, run, and modify Python code to support local experiments and data processing
  • Review transformations and outputs for correctness, quality, and reproducibility
  • Produce clean, reusable, and well-documented workflows suited to benchmark tasks
  • Collaborate on code reviews and task design for real-world AI evaluation work

Requirements

You must have practical, demonstrable experience that matches the skills below. We will evaluate your ability to work with real codebases and produce reproducible data workflows.

  • 3+ years as a data engineer, data scientist, or data-focused software engineer
  • Strong Python ability for data processing, analysis, and workflow validation
  • Experience with structured and unstructured data in real-world codebases
  • Solid machine learning and data-science fundamentals
  • Ability to navigate and modify complex, production-like codebases
  • Ability to review code for maintainability, correctness, and reproducibility
  • Excellent spoken and written English communication skills

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Machine Learning Benchmark Evaluator

Experienced ML engineers wanted for part-time, remote contract work evaluating production-grade model training, evaluation, and inference pipelines; requires 3+ years of ML engineering experience, strong Python, and PyTorch/TensorFlow/JAX familiarity. 20+ hours/week; apply through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +7 more
English
Part-time · Flexible
Expert level

Posted Jul 17, 2026

Machine Learning Engineer, Benchmarking & Evaluation

Join OpenTrain as a remote Machine Learning Engineer focused on benchmark-driven evaluation of real-world ML systems. This contractor role requires 3+ years of ML engineering experience, strong Python skills, and availability 20+ hrs/week with PST overlap.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +7 more
English
Part-time · Flexible
Intermediate level

Posted Jul 16, 2026

AI Benchmark Engineer — Software Engineering

Design and validate multi-agent coding benchmarks using real open-source code changes, Docker, and Python verification scripts. Remote 4-week contractor role for developers in select countries, requiring daily overlap with PST.

Coding & Software
Computer Code Programming
Remote · Bangladesh, Brazil, Colombia +8 more
English
Part-time · Flexible
Entry level

Posted Jul 24, 2026