Skip to content
OpenTrain AIFor AI Companies

Scientific Reasoning Dataset Expert (Physics)

Join OpenTrain AI to design rigorous, reproducible scientific-reasoning datasets that evaluate and improve LLMs. 8-week remote contractor role at 20+ hrs/week requiring 3+ years of physics experience and strong quantitative modeling skills.

OpenTrain AI

Generative AI & RLHF

100% Remote

Worldwide

Eligibility

Entry

Experience

Jul 21, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain (opentrain.ai) is the #1 platform for building careers in AI training and data labeling. OpenTrain AI hires and contracts experts to create high-quality datasets and evaluation tasks that shape how modern AI systems behave.

  • We help people start and grow careers teaching AI through remote, flexible project work.
  • OpenTrain runs rigorous, reproducible projects used by researchers and engineers to evaluate model capabilities.

Why this AI training work matters

AI training (also called data labeling or human feedback work) is the human side of building intelligent systems. Carefully designed evaluation datasets—especially for scientific reasoning—are essential for measuring progress, diagnosing weaknesses, and guiding model improvements.

  • Work is 100% remote and commonly flexible, making it ideal for experienced contributors who want part-time contract work.
  • Contributors directly influence how state-of-the-art language models reason about experiments, data, and scientific laws.

The role: Scientific Reasoning Dataset Expert

You will design and author multi-step scientific reasoning tasks and reproducible evaluation problems that test a model’s ability to analyze data, discover laws, estimate parameters, and predict outcomes. Your work will produce text-based problems, reference solutions, deterministic answers, and scoring rubrics for model assessment.

This is a contractor, part-time assignment reporting to OpenTrain AI for an 8-week engagement. Tasks emphasize scientific accuracy, logical consistency, and reproducibility and require collaboration with reviewers and LLM engineers.

  • Data type: TEXT. Labeling tasks include QUESTION_ANSWERING and EVALUATION_RATING.
  • Employment: CONTRACTOR, PART_TIME. Expected time: 20+ hours/week for 8 weeks.
  • Languages: English. Worldwide applicants accepted.

What you'll do

Create evaluation problems and dataset items that require multi-step scientific reasoning based on experimental, observational, or simulated inputs. Produce clear documentation and scoring guidance so engineers and reviewers can reproduce results and apply consistent ratings.

  • Design scientific reasoning scenarios using experimental, observational, or simulated datasets.
  • Author multi-step tasks that require data analysis, pattern recognition, parameter estimation, and predictive reasoning.
  • Develop evaluation problems focused on law discovery, model selection, consistency checking, and hypothesis validation.
  • Write reference solutions, deterministic answers, and detailed scoring rubrics for model assessment.
  • Collaborate with reviewers and LLM engineers to ensure scientific accuracy, clarity, and reproducibility across datasets.

Requirements

You must meet the stated technical and domain requirements. We preserve every qualification from the listing: strong quantitative skills, rigor, and the ability to communicate complex ideas clearly.

  • Minimum 3+ years of experience in the physics domain.
  • Strong knowledge of scientific reasoning, quantitative modeling, data analysis, and experimental methodologies.
  • Expertise in mathematical modeling, parameter estimation, dimensional analysis, and simulation-based problem solving.
  • Ability to communicate complex scientific concepts clearly through structured documentation.
  • Strong attention to logical consistency, ambiguity control, and reproducible outcomes.
  • Experience level listed as Entry level, with the above 3+ years physics experience required.

Engagement details & next steps

This is a fixed short-term contractor engagement of 8 weeks, remote, and expected to require 20+ hours per week. As a contractor you will not receive paid leave or benefits through the engagement. OpenTrain AI hires and manages the project and reviewer collaboration.

  • Duration: 8 weeks. Time commitment: 20+ hrs/week.
  • Contractor status: freelance engagement; no medical or paid leave provided.
  • To apply, submit your OpenTrain profile and samples that demonstrate physics modeling, reproducible problem design, or prior dataset work.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Physics Research AI Evaluator

Join OpenTrain to evaluate frontier physics research and judge AI model reasoning, earning $80–$110/hr. This part-time, remote US role requires a strong publication record and active research experience (PhD/postdoc preferred).

Generative AI & RLHF
Document
Remote · United States
English
Part-time · Flexible
Expert level
Hourly · $80–$110/hr

Posted Jul 10, 2026

Physics Expert Adjudication Specialist

OpenTrain AI is hiring senior physicists to adjudicate contested physics arguments, compare solutions, and produce rigorous written evaluations used to train next-generation AI systems. This remote contractor role requires a PhD, senior research leadership, ongoing publications, 20+ hours/week, and

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $80–$160/hr

Posted Jul 3, 2026