Skip to content
OpenTrain AIFor AI Companies

Physics Scientific Reasoning Dataset Engineer

Contractor role designing rigorous scientific-reasoning evaluation datasets for advanced AI models; requires 3+ years of physics experience and 20+ hours/week. Work remotely for OpenTrain AI to author multi-step tasks, reference solutions, and scoring rubrics for model assessment.

OpenTrain AI

Generative AI & RLHF

100% Remote

Worldwide

Eligibility

Entry

Experience

Jul 20, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people start and grow careers teaching AI by centralizing specialized AI training work, letting contributors build a unified portfolio, and contracting skilled freelancers directly.

About AI training and why it matters

AI training (data labeling, annotation, and human feedback) is the human side of building modern AI systems. Contributors create and verify the examples models learn from—shaping how state-of-the-art models reason, solve scientific problems, and behave in real-world tasks. Many projects are fully remote and flexible, making this a practical way to work on cutting-edge problems from anywhere.

The role

OpenTrain AI is hiring a contractor to design and create scientific reasoning datasets used to evaluate advanced AI models. This position focuses on building rigorous, reproducible evaluation tasks that probe scientific discovery and multi-step reasoning in physics.

This is a part-time contractor role requiring 20+ hours/week, work in English, and is open worldwide. The work is text-based evaluation design: author problems, reference solutions, and scoring rubrics used for evaluation ratings.

What you'll do

  • Design scientific reasoning scenarios that use experimental, observational, or simulated datasets.
  • Author multi-step tasks that require data analysis, pattern recognition, parameter estimation, and predictive reasoning.
  • Develop evaluation problems focused on law discovery, model selection, consistency checking, and hypothesis validation.
  • Create reference solutions, deterministic answers, and detailed scoring rubrics for model assessment.
  • Collaborate with reviewers and LLM engineers to ensure scientific accuracy, clarity, and reproducibility across datasets.

Requirements

  • 3+ years of professional experience in physics or a closely related scientific field.
  • Strong knowledge of scientific reasoning, quantitative modeling, data analysis, and experimental methodologies.
  • Expertise in mathematical modeling, parameter estimation, dimensional analysis, and simulation-based problem solving.
  • Ability to communicate complex scientific concepts, assumptions, and solutions clearly through structured documentation.
  • Strong attention to logical consistency, ambiguity control, and producing reproducible outcomes.

Who should apply

This role is for practitioners with a strong physics background who enjoy turning scientific problems into clear, testable evaluation tasks. The listing notes an entry-level experience level, while the role explicitly requires 3+ years of professional physics experience—applicants should meet the stated experience and skills.

  • You enjoy designing reproducible experiments, writing clear problem statements, and producing precise grading rubrics.
  • You are comfortable working asynchronously with engineers and reviewers to iterate on dataset quality.
  • Fluent English for written documentation and rubric creation.

How it works

OpenTrain AI is the hiring and contracting organization for this role. Successful applicants will be contracted to produce text-based evaluation materials (data type: TEXT) used for evaluation ratings (label type: EVALUATION_RATING).

Apply through your OpenTrain profile, and, if selected, you will collaborate with reviewers and engineers to deliver tasks, solutions, and rubrics. The position is remote and flexible but requires a commitment of 20+ hours per week.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Scientific Reasoning Dataset Expert (Physics)

Join OpenTrain AI to design rigorous, reproducible scientific-reasoning datasets that evaluate and improve LLMs. 8-week remote contractor role at 20+ hrs/week requiring 3+ years of physics experience and strong quantitative modeling skills.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 21, 2026

Physics Reasoning Evaluator (BS/MS/PhD Required)

Join OpenTrain AI to evaluate and improve physics-focused AI outputs—paid contract work at $80/hr, part-time (minimum ~17–20 hrs/week). Use your physics degree to spot subtle errors, write step-by-step solutions, and rate model responses with detailed rubrics.

Generative AI & RLHF
Video
Remote · Worldwide
Part-time · Flexible
Entry level
Hourly · $80/hr

Posted Oct 24, 2025

Chemical Reasoning & Discovery Engineer

Join OpenTrain as a remote Chemical Reasoning & Discovery Engineer creating chemistry-focused reasoning datasets and reproducible evaluation rubrics for LLMs. This 8-week, part-time contractor role (~20+ hrs/week) requires 4 hours overlap with PST and is ideal for experienced chemists.

Generative AI & RLHF
Document
Remote · Worldwide
English
Part-time · Flexible
Intermediate level

Posted Jul 17, 2026