Skip to content
OpenTrain AIFor AI Companies

Data Science AI Evaluation Expert

Design enterprise-grade data science scenarios, reference analyses, and evaluation rubrics that teach AI to reason like a senior data leader. This fully remote freelance contract pays $60–$70 per hour for a 40-hour weekly commitment.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $60–$70/hr

$60–$70/hr

Compensation

Worldwide

Eligibility

Intermediate

Experience

Aug 7, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the #1 platform for finding and building careers in AI training and data labeling. We help contributors discover specialized projects, build a professional profile, and grow their experience teaching and evaluating advanced AI systems.

Creating an OpenTrain account is free, and this opportunity is managed as a freelance contracting engagement with OpenTrain AI.

About AI Evaluation Work

AI training is the human side of building artificial intelligence. Experts create examples, assess model outputs, and define evaluation standards that help AI systems become more accurate, useful, and reliable.

In this role, your enterprise data science judgment will help evaluate whether AI systems can reason through the scale, complexity, governance requirements, and business-critical decisions found in large organizations.

The Role

OpenTrain AI is seeking a Data Science AI Evaluation Expert to build evaluation tasks for AI systems operating in Fortune 500 enterprise data and analytics contexts. You will design realistic data science scenarios, draft reference outputs, and create rubrics that distinguish senior-level judgment from generic textbook knowledge.

Assignments will reflect large-scale data operations, complex model development, enterprise infrastructure, cross-functional governance, and the practical stakes of business-critical analytics.

  • Contract type: Freelance contractor
  • Pay: $60–$70 per hour
  • Expected commitment: 40 hours per week
  • Work arrangement: Fully remote and worldwide
  • Language: English
  • Employment type: Part-time contract opportunity, not a permanent employee role

What You'll Do

You will translate hands-on experience with enterprise data science into challenging AI evaluation content. Your work will cover the technical, methodological, and organizational decisions that senior data and analytics leaders make in production environments.

  • Construct enterprise data science scenarios involving large-scale predictive modeling, multi-stakeholder analytics governance, and complex data infrastructure decisions.
  • Build tasks across machine learning model development, enterprise data pipelines, business intelligence at scale, experimentation and causal inference, and data strategy.
  • Develop data and MLOps scenarios using tools such as Snowflake, Databricks, Python, R, SQL, Tableau, Power BI, SageMaker, Vertex AI, and MLflow.
  • Apply statistical rigor, A/B testing frameworks, model validation practices, and MLOps best practices.
  • Produce reference analyses, model documentation, and executive-level insights.
  • Author rubrics that identify authentic enterprise data science judgment rather than tutorial-level recall.

Required Experience

This opportunity is intended for an experienced data science, analytics, or machine learning professional with direct responsibility for enterprise systems and initiatives. You should be comfortable evaluating nuanced technical decisions in environments where governance, privacy, scale, and stakeholder alignment matter.

  • 5+ years as a data scientist, analytics leader, or ML engineer at a Fortune 500 technology or enterprise organization, or within a Fortune 500 data and analytics organization.
  • Examples of relevant organizations include Google, Meta, Amazon, Microsoft, Netflix, JPMorgan, UPS, Unilever, PepsiCo, and Walmart.
  • Direct ownership of Fortune 500 data products, analytics initiatives, or machine learning systems in production.
  • Fluency in enterprise data science tooling and methodologies.
  • Understanding of Fortune 500 data governance, privacy compliance, and cross-functional stakeholder alignment.

Helpful Background and Application Details

Prior experience authoring evaluation rubrics, technical curricula, or model documentation is helpful. The work is especially suited to professionals who can explain complex data science decisions clearly and recognize the difference between surface-level answers and sound enterprise practice.

This is a remote freelance contracting opportunity. The structured project requirement is 20 or more hours per week, while the stated role commitment is 40 hours per week.

  • Relevant backgrounds: data science, analytics leadership, machine learning engineering, enterprise data platforms, or MLOps.
  • Relevant output types: text generation, evaluation, reference analyses, and rating rubrics.
  • Compensation is paid hourly within the listed $60–$70 range.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Data Science AI Evaluation Expert

Use your data science expertise to evaluate, fact-check, and improve AI-generated content and analytical outputs. This remote, part-time contractor role offers $100–$200 per hour and requires 20+ hours weekly.

Generative AI & RLHF
Document
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $100–$200/hr

Posted Aug 4, 2026

Data Science AI Evaluation Expert

Join OpenTrain to evaluate AI-generated and human-created data science work: design grading criteria, score technical deliverables, and write defensible evaluations. Remote, contract role for experienced data scientists with 20+ hours/week availability and $100–$150/hr pay.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Intermediate level
Hourly · $100–$150/hr

Posted Jul 29, 2026

AI Evaluation Benchmark Researcher

Design and author multi-step scientific evaluation tasks for frontier AI models in a full-time remote US contractor role paying $60–$90/hr. Expect ~35 hours/week building Python reference solutions, defining rigorous criteria, and reviewing model attempts.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Entry level
Hourly · $60–$90/hr

Posted Jul 29, 2026