SWE Bench Data Engineer / Data Scientist — Benchmark Evaluation
Join OpenTrain AI as a contract, part-time SWE Bench contributor working 20+ hours/week on benchmark-driven data engineering and data-science evaluation. Use Python to build reproducible pipelines, process structured and unstructured data, and validate real-world workflows (English required).
Coding & Software
Remote
10 countries
Eligibility
Intermediate
Experience
Jul 16, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Brazil Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the centralized, open platform where people build careers in AI training and data labeling. Contributors use OpenTrain to find projects, consolidate work history, and develop a reusable portfolio that demonstrates their skills across AI evaluation and annotation work.
OpenTrain AI is the hiring and contracting organization for this role. Creating an OpenTrain account is free and lets you apply, manage opportunities, and build a record of benchmark and annotation work that employers respect.
A focused career platform for people who train and evaluate AI systems
Work is remote, flexible, and directly shapes how modern AI behaves
Free to join and build a unified portfolio across projects
About AI Training Work
AI training (data labeling, annotation, and human evaluation) is the human work behind modern AI systems. Contributors annotate, validate, and review data and model outputs so researchers and engineers can measure and improve model behavior.
This role sits at the intersection of data engineering and data science: you will support benchmark-driven evaluations that mimic production workflows to ensure results are correct, reproducible, and useful for model development.
100% remote, often flexible hours and part-time-friendly
Work ranges from beginner-friendly annotation to advanced engineering and evaluation
Contributors directly affect how state-of-the-art AI systems are assessed
The Role
We are hiring experienced SWE Bench Data Engineers / Data Scientists to perform benchmark-driven evaluation work focused on real-world data engineering and data science workflows. The work centers on structured and unstructured datasets, production-like data pipelines, feature preparation, data validation, and Python-based experimentation in complex codebases.
This is a contract, part-time role expecting 20+ hours per week. Candidates must be fluent in spoken and written English and be located in one of the approved countries: IN, PK, NG, KE, EG, GH, BD, TR, BR, MX. Experience level: Intermediate (3+ years).
Build, validate, and maintain data pipelines for benchmarking and evaluation workflows
Process, analyze, and prepare structured and unstructured data for data-science use cases
Write, run, and modify Python code to support local experiments and data processing
Review transformations and outputs for correctness, quality, and reproducibility
Produce clean, reusable, and well-documented workflows suited to benchmark tasks
Collaborate on code reviews and task design for real-world AI evaluation work
Requirements
You must have practical, demonstrable experience that matches the skills below. We will evaluate your ability to work with real codebases and produce reproducible data workflows.
3+ years as a data engineer, data scientist, or data-focused software engineer
Strong Python ability for data processing, analysis, and workflow validation
Experience with structured and unstructured data in real-world codebases
Solid machine learning and data-science fundamentals
Ability to navigate and modify complex, production-like codebases
Ability to review code for maintainability, correctness, and reproducibility
Excellent spoken and written English communication skills
Experienced ML engineers wanted for part-time, remote contract work evaluating production-grade model training, evaluation, and inference pipelines; requires 3+ years of ML engineering experience, strong Python, and PyTorch/TensorFlow/JAX familiarity. 20+ hours/week; apply through OpenTrain.
Join OpenTrain as a remote Machine Learning Engineer focused on benchmark-driven evaluation of real-world ML systems. This contractor role requires 3+ years of ML engineering experience, strong Python skills, and availability 20+ hrs/week with PST overlap.
Design and validate multi-agent coding benchmarks using real open-source code changes, Docker, and Python verification scripts. Remote 4-week contractor role for developers in select countries, requiring daily overlap with PST.