Design enterprise-grade data science scenarios, reference analyses, and evaluation rubrics that teach AI to reason like a senior data leader. This fully remote freelance contract pays $60–$70 per hour for a 40-hour weekly commitment.
Generative AI & RLHF
100% Remote Hourly · $60–$70/hr
$60–$70/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Aug 7, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the #1 platform for finding and building careers in AI training and data labeling. We help contributors discover specialized projects, build a professional profile, and grow their experience teaching and evaluating advanced AI systems.
Creating an OpenTrain account is free, and this opportunity is managed as a freelance contracting engagement with OpenTrain AI.
About AI Evaluation Work
AI training is the human side of building artificial intelligence. Experts create examples, assess model outputs, and define evaluation standards that help AI systems become more accurate, useful, and reliable.
In this role, your enterprise data science judgment will help evaluate whether AI systems can reason through the scale, complexity, governance requirements, and business-critical decisions found in large organizations.
The Role
OpenTrain AI is seeking a Data Science AI Evaluation Expert to build evaluation tasks for AI systems operating in Fortune 500 enterprise data and analytics contexts. You will design realistic data science scenarios, draft reference outputs, and create rubrics that distinguish senior-level judgment from generic textbook knowledge.
Assignments will reflect large-scale data operations, complex model development, enterprise infrastructure, cross-functional governance, and the practical stakes of business-critical analytics.
Contract type: Freelance contractor
Pay: $60–$70 per hour
Expected commitment: 40 hours per week
Work arrangement: Fully remote and worldwide
Language: English
Employment type: Part-time contract opportunity, not a permanent employee role
What You'll Do
You will translate hands-on experience with enterprise data science into challenging AI evaluation content. Your work will cover the technical, methodological, and organizational decisions that senior data and analytics leaders make in production environments.
Construct enterprise data science scenarios involving large-scale predictive modeling, multi-stakeholder analytics governance, and complex data infrastructure decisions.
Build tasks across machine learning model development, enterprise data pipelines, business intelligence at scale, experimentation and causal inference, and data strategy.
Develop data and MLOps scenarios using tools such as Snowflake, Databricks, Python, R, SQL, Tableau, Power BI, SageMaker, Vertex AI, and MLflow.
Apply statistical rigor, A/B testing frameworks, model validation practices, and MLOps best practices.
Produce reference analyses, model documentation, and executive-level insights.
Author rubrics that identify authentic enterprise data science judgment rather than tutorial-level recall.
Required Experience
This opportunity is intended for an experienced data science, analytics, or machine learning professional with direct responsibility for enterprise systems and initiatives. You should be comfortable evaluating nuanced technical decisions in environments where governance, privacy, scale, and stakeholder alignment matter.
5+ years as a data scientist, analytics leader, or ML engineer at a Fortune 500 technology or enterprise organization, or within a Fortune 500 data and analytics organization.
Examples of relevant organizations include Google, Meta, Amazon, Microsoft, Netflix, JPMorgan, UPS, Unilever, PepsiCo, and Walmart.
Direct ownership of Fortune 500 data products, analytics initiatives, or machine learning systems in production.
Fluency in enterprise data science tooling and methodologies.
Understanding of Fortune 500 data governance, privacy compliance, and cross-functional stakeholder alignment.
Helpful Background and Application Details
Prior experience authoring evaluation rubrics, technical curricula, or model documentation is helpful. The work is especially suited to professionals who can explain complex data science decisions clearly and recognize the difference between surface-level answers and sound enterprise practice.
This is a remote freelance contracting opportunity. The structured project requirement is 20 or more hours per week, while the stated role commitment is 40 hours per week.
Relevant backgrounds: data science, analytics leadership, machine learning engineering, enterprise data platforms, or MLOps.
Relevant output types: text generation, evaluation, reference analyses, and rating rubrics.
Compensation is paid hourly within the listed $60–$70 range.
Use your data science expertise to evaluate, fact-check, and improve AI-generated content and analytical outputs. This remote, part-time contractor role offers $100–$200 per hour and requires 20+ hours weekly.
Join OpenTrain to evaluate AI-generated and human-created data science work: design grading criteria, score technical deliverables, and write defensible evaluations. Remote, contract role for experienced data scientists with 20+ hours/week availability and $100–$150/hr pay.
Design and author multi-step scientific evaluation tasks for frontier AI models in a full-time remote US contractor role paying $60–$90/hr. Expect ~35 hours/week building Python reference solutions, defining rigorous criteria, and reviewing model attempts.