Help evaluate AI systems on real-world data science tasks, including modeling, experimentation, and technical reports. This remote contractor role pays $100-$150 per hour and requires 20+ hours weekly.
About OpenTrain
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and apply in minutes.
Creating an OpenTrain account is free, and this opportunity offers a way to contribute specialized data science expertise to the rapidly growing AI training industry.
About AI Evaluation Work
AI training is the human side of building artificial intelligence. Expert contributors assess model outputs, create evaluation standards, and provide clear feedback that helps AI systems perform more reliably in real-world settings.
In this role, your data science judgment will help determine whether AI-generated work meets rigorous standards for analysis, statistical reasoning, machine learning, experimentation, and technical communication.
- Remote work with flexible contractor scheduling
- Work at the intersection of data science and cutting-edge AI development
- Use evidence-based evaluation to improve the quality of AI systems
The Role
OpenTrain AI is recruiting a Data Science AI Evaluation Expert for a talent network supporting future projects. You will evaluate how effectively AI systems perform real-world data science work and assess both AI-generated and human-created deliverables.
This is a remote, hourly contractor role paying $100-$150 per hour. The expected time requirement is 20 or more hours per week, with a default commitment of 40 hours per week.
- Role type: Part-time contractor
- Pay: $100-$150 per hour
- Workload: 20+ hours per week, with a default 40-hour commitment
- Working language: English
What You’ll Do
You will create and apply precise evaluation standards for complex data science work. Your assessments should be consistent, defensible, and supported by detailed written reasoning.
You will also incorporate structured feedback from senior reviewers and refine submitted work as evaluation standards develop.
- Design task-specific grading criteria for exploratory data analyses
- Evaluate statistical modeling work, machine learning pipelines, and feature engineering
- Assess experimentation and A/B test write-ups, including causal inference considerations
- Review technical reports and notebooks
- Evaluate AI-generated or human-created data science deliverables
- Provide detailed written justifications for scores and judgments
- Apply consistent, evidence-based standards so assessments are reproducible
- Incorporate senior reviewer feedback and iterate on submitted work
Requirements
This opportunity is intended for candidates with professional data science experience and a strong record of working with technically complex deliverables. The listing identifies the experience level as entry level, while also requiring at least one year of professional experience and background at a leading technology, research, or quantitative firm.
- At least 1 year of professional data science experience
- Experience at a leading technology, research, or quantitative firm, such as a top FAANG company, AI lab, top-tier quantitative fund, or equivalent
- Strong command of Python and SQL
- Strong command of statistical modeling and machine learning
- Experience with experimentation and causal inference
- Exceptional written communication for explaining technical findings clearly
- Detail-oriented and consistent approach to evaluating complex work
- Comfort receiving feedback and calibrating judgment against established standards
Who Should Apply
This role may suit data science professionals who enjoy carefully reviewing analytical work and explaining why an approach is correct, incomplete, or flawed. It is especially relevant for candidates comfortable evaluating notebooks, models, experiments, and technical reports against explicit standards.
Strong writing matters because every score must be supported by a clear, technically accurate explanation. The work also calls for consistency when applying judgment across varied data science submissions.
- Data scientists with experience in statistical or machine learning projects
- Professionals who can assess experimental design and causal reasoning
- Technical reviewers who communicate findings with precision
- Candidates comfortable receiving feedback and improving evaluation decisions
Location and How to Apply
This role is remote and available to candidates located in Austria, Belgium, Bulgaria, Canada, Cyprus, Czechia, Germany, Denmark, Estonia, Spain, Finland, France, the United Kingdom, Greece, Croatia, Hungary, Ireland, Italy, Lithuania, Luxembourg, Latvia, Malta, the Netherlands, Poland, Portugal, Romania, Sweden, Slovenia, Slovakia, or the United States.
Apply through OpenTrain AI to be considered for this contractor talent network. OpenTrain helps contributors start and grow careers in AI training and data labeling by connecting their skills with projects shaping how modern AI systems work.
- Remote opportunity
- Eligible locations are limited to the countries listed above
- Apply in English through OpenTrain AI
- Create an OpenTrain account for free