Join OpenTrain as a contractor designing and validating data pipelines and benchmark evaluation tasks for AI systems. Work part-time (20+ hrs/wk) with Python on production-like datasets; applicants need 3+ years in data engineering, data science, or data-focused software engineering.
Generative AI & RLHF
Remote
10 countries
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Brazil Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people building careers in AI training and data labeling. We connect contractors with real evaluation and annotation work, help you build a unified portfolio, and support durable freelance careers in the human side of AI.
For this role, OpenTrain is the hiring and contracting organization: you'll join a distributed team working on benchmark-driven evaluation tasks that shape how advanced AI systems are measured and improved.
Why AI training and evaluation work matters
AI training (data labeling, annotation, and evaluation) is the human foundation of modern AI. Contributors prepare, process, and validate the examples models learn from — a fast-growing way to work remotely with flexible hours and direct impact on model behavior.
This role sits at the intersection of data engineering and model evaluation: you'll build reproducible data workflows and tasks used to benchmark model performance in realistic settings.
The role
OpenTrain is hiring a Data Engineering & Data Science AI Evaluation Engineer to design, build, and validate data pipelines and evaluation tasks used in benchmarking advanced AI models. This is a hands-on contractor role working with production-like structured and unstructured datasets and Python code.
You will collaborate with researchers and engineers to create rigorous, reproducible evaluation workflows and participate in code reviews to maintain quality.
Contract type: Contractor, part-time.
Minimum commitment: 20+ hours/week with at least 4 hours/day and 4 hours daily overlap with Pacific Time (PST).
Initial duration: 3-month contract (adjustable based on engagement).
What you'll do
Deliver clean, documented data engineering work that supports benchmark-style evaluation tasks and experiments. You will write and run Python code locally, prepare features, and ensure transformations are correct and reproducible.
Work with structured and unstructured datasets to support SWE Bench-style evaluation tasks.
Design, build, and validate data pipelines used in benchmarking and evaluation workflows.
Perform data processing, analysis, feature preparation, and validation for data science use cases.
Write, run, and modify Python code to process data and support experiments locally.
Evaluate data quality, transformations, and outputs for correctness and reproducibility.
Create clean, well-documented, and reusable data workflows suitable for benchmarking.
Participate in code reviews to ensure high standards of code quality and maintainability.
Collaborate with researchers and engineers to design challenging, real-world data engineering and data science tasks for AI systems.
Requirements
Candidates must meet the stated technical experience and availability requirements. We will evaluate your ability to work with complex codebases, write maintainable Python, and reason about data and model-related workflows.
Minimum 3+ years experience as a Data Engineer, Data Scientist, or data-focused Software Engineer.
Strong proficiency in Python for data engineering and data science workflows.
Demonstrable experience with data processing, analysis, and model-related workflows.
Solid understanding of machine learning and data science fundamentals.
Experience working with both structured and unstructured data.
Ability to understand, navigate, and modify complex, real-world codebases.
Experience writing readable, reusable, maintainable, and well-documented code.
Strong problem-solving skills, including experience with algorithmic or data-intensive problems.
Excellent spoken and written English communication skills.
Helpful background
The following are nice-to-have but not required. They indicate familiarity with benchmark-driven evaluation and designing tasks that stress real-world model behavior.
Experience with SWE Bench or similar benchmark-driven evaluation projects.
Background in designing AI model evaluation tasks or benchmarks.
How to apply and next steps
Apply through your OpenTrain account and include examples of relevant Python work, data pipelines, or evaluation projects. Highlight projects demonstrating reproducible workflows and your role in design or validation.
Because this role focuses on benchmarking code and datasets, applications that show concrete repositories, scripts, or brief write-ups of past evaluation work will help your candidacy stand out.
Languages: English required.
Eligible countries: IN, PK, NG, KE, EG, GH, BD, TR, BR, MX (applications accepted from these locations).
Data types and label work: this role focuses on COMPUTER_CODE_PROGRAMMING data and label tasks like EVALUATION_RATING and COMPUTER_PROGRAMMING_CODING.
Join OpenTrain AI to evaluate and improve AI-driven infrastructure automation: generate prompts, rate DevOps runbooks, and assess self-hosted platforms. Part-time contractor role requiring 2+ years DevOps experience, annotation/Q A skills, English B2+, 20+ hrs/week, $15–$45/hr.
Join OpenTrain to evaluate and improve AI-generated civil engineering outputs, providing technical feedback on calculations, assumptions, safety, and professional judgment. Remote contractor role, 20+ hours/week, paid $45–$100/hr; English fluency and 4+ years civil engineering experience required.
Join OpenTrain as an AI Model Evaluation Developer to write and maintain code, run model benchmarks, rank responses, and build datasets for fine-tuning and RLHF. This remote, part-time contractor role requires strong Python and JavaScript/TypeScript skills and 20+ hours/week.