OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring and contracting for this role, connecting skilled contributors with hands-on projects that help shape how advanced AI systems are built and evaluated.
Creating an OpenTrain account is free, and applicants can apply in minutes while building a professional profile focused on AI training and data work.
About AI Evaluation Work
AI evaluation is the human side of improving artificial intelligence. Engineers and data specialists prepare datasets, create benchmarks, review outputs, and validate workflows so developers can measure whether AI systems perform accurately and reliably.
This work brings together software engineering, data science, and machine learning in a fast-growing field. Contributors work with production-like problems and help shape the behavior and capabilities of cutting-edge AI systems.
The Role
OpenTrain is recruiting a Data Engineering and Data Science AI Evaluation Engineer to design and validate data pipelines and evaluation tasks used to benchmark advanced AI systems. This is a hands-on contractor position involving production-like datasets, Python code, and real-world data workflows.
You will help create challenging, realistic tasks for AI systems and evaluate the quality, correctness, and reproducibility of the workflows and outputs used to measure model performance.
Contractor and part-time engagement
Three-month contract, adjustable based on engagement
At least 4 hours per day and 20 hours per week
Four hours of daily overlap with Pacific Standard Time
Available to candidates in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Türkiye, Brazil, and Mexico
English language communication required
What You'll Do
You will work across data engineering and data science workflows, using Python and complex real-world codebases to build reliable benchmarking and evaluation tasks. The role combines implementation, analysis, validation, documentation, and collaboration with researchers and engineers.
Work with structured and unstructured datasets for SWE Bench-style evaluation tasks
Design, build, and validate data pipelines for benchmarking and evaluation workflows
Perform data processing, analysis, feature preparation, and validation for data science use cases
Write, run, and modify Python code locally to process data and support experiments
Evaluate data quality, transformations, and outputs for correctness and reproducibility
Create clean, documented, and reusable workflows suitable for benchmarking
Participate in code reviews focused on quality and maintainability
Collaborate with researchers and engineers to design challenging, real-world AI tasks
Requirements
This role requires at least three years of overall experience as a Data Engineer, Data Scientist, or data-focused Software Engineer. You should be comfortable applying data engineering and data science methods in Python and working independently with complex technical problems.
At least 3 years of experience as a Data Engineer, Data Scientist, or data-focused Software Engineer
Strong proficiency in Python for data engineering and data science workflows
Demonstrable experience with data processing, analysis, and model-related workflows
Solid understanding of machine learning and data science fundamentals
Experience with structured and unstructured data
Ability to understand, navigate, and modify complex, real-world codebases
Experience writing readable, reusable, maintainable, and well-documented code
Strong problem-solving skills with algorithmic or data-intensive problems
Excellent spoken and written English communication skills
Helpful Background
Experience with benchmark-driven evaluation can help you contribute quickly, particularly when designing tasks that reflect realistic engineering and data science challenges.
Experience with SWE Bench or similar benchmark-driven evaluation projects
Background designing AI model evaluation tasks
Why Work in AI Training
AI training and data labeling work is a growing way to work in technology without needing to build an entire AI system alone. Specialists contribute directly by preparing examples, testing model behavior, evaluating results, and improving the datasets and workflows behind modern AI.
Remote work with a computer and internet connection
Flexible part-time opportunities that can fit around other commitments
Direct involvement with state-of-the-art AI systems
Opportunities to build experience in an expanding technical field
How to Apply
Create a free OpenTrain account, build your profile around your data engineering, data science, Python, and evaluation experience, and apply through OpenTrain. Be prepared to demonstrate your ability to work with complex codebases, production-like data, and reproducible evaluation workflows.
Evaluate AI-powered developer workflows through hands-on coding environments, GitHub, CI/CD, and technical assessment. This worldwide contractor role pays $50-$70 per hour for 20+ hours weekly.
Use your software engineering judgment to evaluate AI-generated technical content, refine prompts, fact-check claims, and create high-quality engineering artifacts. Work remotely as a flexible contractor for 20+ hours per week.
Use your frontend engineering expertise to evaluate code, architecture, and technical decisions that help improve AI systems. This remote contractor role offers $90-$140 per hour and requires 20+ hours each week.