Join OpenTrain AI to design rigorous, reproducible scientific-reasoning datasets that evaluate and improve LLMs. 8-week remote contractor role at 20+ hrs/week requiring 3+ years of physics experience and strong quantitative modeling skills.
Generative AI & RLHF
100% Remote
Worldwide
Eligibility
Entry
Experience
Jul 21, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain (opentrain.ai) is the #1 platform for building careers in AI training and data labeling. OpenTrain AI hires and contracts experts to create high-quality datasets and evaluation tasks that shape how modern AI systems behave.
We help people start and grow careers teaching AI through remote, flexible project work.
OpenTrain runs rigorous, reproducible projects used by researchers and engineers to evaluate model capabilities.
Why this AI training work matters
AI training (also called data labeling or human feedback work) is the human side of building intelligent systems. Carefully designed evaluation datasets—especially for scientific reasoning—are essential for measuring progress, diagnosing weaknesses, and guiding model improvements.
Work is 100% remote and commonly flexible, making it ideal for experienced contributors who want part-time contract work.
Contributors directly influence how state-of-the-art language models reason about experiments, data, and scientific laws.
The role: Scientific Reasoning Dataset Expert
You will design and author multi-step scientific reasoning tasks and reproducible evaluation problems that test a model’s ability to analyze data, discover laws, estimate parameters, and predict outcomes. Your work will produce text-based problems, reference solutions, deterministic answers, and scoring rubrics for model assessment.
This is a contractor, part-time assignment reporting to OpenTrain AI for an 8-week engagement. Tasks emphasize scientific accuracy, logical consistency, and reproducibility and require collaboration with reviewers and LLM engineers.
Data type: TEXT. Labeling tasks include QUESTION_ANSWERING and EVALUATION_RATING.
Employment: CONTRACTOR, PART_TIME. Expected time: 20+ hours/week for 8 weeks.
Create evaluation problems and dataset items that require multi-step scientific reasoning based on experimental, observational, or simulated inputs. Produce clear documentation and scoring guidance so engineers and reviewers can reproduce results and apply consistent ratings.
Design scientific reasoning scenarios using experimental, observational, or simulated datasets.
Author multi-step tasks that require data analysis, pattern recognition, parameter estimation, and predictive reasoning.
Develop evaluation problems focused on law discovery, model selection, consistency checking, and hypothesis validation.
Write reference solutions, deterministic answers, and detailed scoring rubrics for model assessment.
Collaborate with reviewers and LLM engineers to ensure scientific accuracy, clarity, and reproducibility across datasets.
Requirements
You must meet the stated technical and domain requirements. We preserve every qualification from the listing: strong quantitative skills, rigor, and the ability to communicate complex ideas clearly.
Minimum 3+ years of experience in the physics domain.
Strong knowledge of scientific reasoning, quantitative modeling, data analysis, and experimental methodologies.
Expertise in mathematical modeling, parameter estimation, dimensional analysis, and simulation-based problem solving.
Ability to communicate complex scientific concepts clearly through structured documentation.
Strong attention to logical consistency, ambiguity control, and reproducible outcomes.
Experience level listed as Entry level, with the above 3+ years physics experience required.
Engagement details & next steps
This is a fixed short-term contractor engagement of 8 weeks, remote, and expected to require 20+ hours per week. As a contractor you will not receive paid leave or benefits through the engagement. OpenTrain AI hires and manages the project and reviewer collaboration.
Duration: 8 weeks. Time commitment: 20+ hrs/week.
Contractor status: freelance engagement; no medical or paid leave provided.
To apply, submit your OpenTrain profile and samples that demonstrate physics modeling, reproducible problem design, or prior dataset work.
Join OpenTrain AI to evaluate and improve physics-focused AI outputs—paid contract work at $80/hr, part-time (minimum ~17–20 hrs/week). Use your physics degree to spot subtle errors, write step-by-step solutions, and rate model responses with detailed rubrics.
Join OpenTrain to evaluate frontier physics research and judge AI model reasoning, earning $80–$110/hr. This part-time, remote US role requires a strong publication record and active research experience (PhD/postdoc preferred).
OpenTrain AI is hiring senior physicists to adjudicate contested physics arguments, compare solutions, and produce rigorous written evaluations used to train next-generation AI systems. This remote contractor role requires a PhD, senior research leadership, ongoing publications, 20+ hours/week, and