Design and solve advanced physics problems to probe large language models' multi-step reasoning and symbolic skills; remote, contract, 20+ hours/week working with OpenTrain AI to build benchmarks from undergraduate to PhD levels.
Generative AI & RLHF
100% Remote
Worldwide
Eligibility
Intermediate
Experience
Jul 17, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people who build careers in AI training and data labeling. We help contributors discover projects, consolidate experience, and grow a durable freelance career teaching AI — all from one profile.
OpenTrain AI is the hiring and contracting organization for this role. We focus on remote, flexible work that lets subject-matter experts contribute directly to how state-of-the-art systems learn.
About AI training and why it matters
AI training (data labeling / annotation) is the human side of building intelligent systems: people create, evaluate, and refine examples that models learn from. This work is highly flexible, often entry-accessible, and places contributors on the cutting edge of AI behavior.
In this role you'll help shape how language models understand and solve physics problems — a critical step toward reliable, explainable scientific reasoning in AI.
The role
We're hiring a remote Physics LLM Evaluation Expert to design challenging physics problems, write clear step-by-step solutions, and help define evaluation benchmarks spanning early undergraduate through PhD-level topics.
This is a contract, part-time role with a time expectation of 20+ hours per week. Work is performed remotely and independently while collaborating with LLM researchers and the OpenTrain team.
Employment type: Contractor, Part-time
Time commitment: 20+ hours/week
Work type: Remote, worldwide (English required)
What you'll do
Your core responsibility is to design problems that reveal models' weaknesses and produce high-quality, structured solutions that demonstrate correct reasoning and common failure modes.
Create original physics problems across mechanics, E&M, optics, thermodynamics, and statistical physics.
Write clear, step-by-step solutions with structured reasoning and careful explanations.
Design items that probe abstraction, multi-step reasoning, and symbolic manipulation in LLMs.
Collaborate with LLM researchers to align tasks with evaluation goals and benchmark criteria.
Help shape evaluation benchmarks mapped to curricula from undergraduate to PhD levels.
Requirements
Candidates must show strong physics fundamentals, excellent written communication in English, and the ability to produce clear, logical solutions independently in a remote setting.
Strong foundation in physics including mechanics, electromagnetism, optics, thermodynamics, and statistical physics.
Ability to design original problems and produce clear, step-by-step solutions.
Excellent structured communication and collaboration skills for remote work.
Strong analytical thinking and research skills; self-motivated and able to work independently.
Helpful background
The role favors candidates with advanced study or prior experience creating detailed annotations, evaluations, or instructional solutions, but relevant professional experience is also valuable.
Master’s, Ph.D., or postdoctoral study in Physics, Applied Physics, or a closely related field is helpful.
Experience creating detailed annotations, constructive feedback, or educational problem sets is a plus.
Comfort solving complex problems with a logical, step-by-step approach.
How the work is done and how to apply
You will work on text-based evaluation tasks: generating physics problems and solutions and rating model outputs (label types: TEXT_GENERATION and EVALUATION_RATING). Collaboration is remote with regular alignment to evaluation goals.
To apply, create an OpenTrain account and submit your profile and relevant examples of problem design or solution write-ups. Successful applicants will be onboarded and briefed on task guidelines and benchmark objectives.
Data type: Text; primary label types TEXT_GENERATION and EVALUATION_RATING.
Language: English required.
Worldwide applicants accepted; work performed remotely.
Application steps: create an OpenTrain profile, submit examples of physics problems/solutions, and complete onboarding guidelines.
Join OpenTrain to design and solve advanced physics problems and build evaluation benchmarks that fine-tune large language models. This remote, part-time contractor role (20+ hrs/week) suits PhD-level physicists or equivalent with strong symbolic and multi-step reasoning skills.
Join OpenTrain to design and solve advanced physics problems that probe LLM reasoning and symbolic skills; remote, part-time contractor work (20+ hrs/week) for candidates with graduate-level physics experience and strong written English.
Design and solve challenging physics problems and write step-by-step solutions to probe and evaluate large language models across undergraduate to PhD topics. Remote, part-time contract (~20+ hrs/week); requires strong physics foundations and Python scientific-computing skills.