Design and solve challenging physics problems and write step-by-step solutions to probe and evaluate large language models across undergraduate to PhD topics. Remote, part-time contract (~20+ hrs/week); requires strong physics foundations and Python scientific-computing skills.
Generative AI & RLHF
100% Remote
Worldwide
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people building careers in AI training and data labeling. We help contributors discover specialized projects, build a unified AI-training portfolio, and grow a durable freelance career doing hands-on work that shapes how AI systems behave.
Why AI training matters
AI training (also called data labeling or human feedback work) is the human side of building modern AI: people create examples, evaluate model outputs, and provide the structured reasoning that helps models learn. This role puts you on the cutting edge, directly influencing how generative models reason about challenging physics problems.
Work is 100% remote and often flexible — great for part-time contributors.
You will directly shape evaluation benchmarks that guide model improvements.
The role
As a Physics LLM Evaluation Specialist you will design difficult physics problems that reveal LLM strengths and weaknesses, produce clear step-by-step solutions with verifiable numerical results, and collaborate with LLM researchers to align tasks with evaluation goals. This is hands-on evaluation work where precise reasoning, structured explanations, and reproducible outcomes matter.
Contract, part-time engagement (20+ hours/week).
Work focuses on text inputs/outputs: problem prompts, model responses, and evaluation annotations.
What you'll do
Create physics problems across a range of topics and academic levels (early undergraduate through PhD) that test abstraction, multi-step reasoning, and symbolic manipulation.
Write rigorous, step-by-step solutions with clear reasoning and verifiable numeric answers.
Use Python and scientific libraries to compute, verify, and illustrate solutions when needed.
Collaborate with LLM researchers to refine evaluation goals and benchmark design.
Help define and iterate physics evaluation benchmarks and rubrics.
Requirements
Strong physics foundation across undergraduate to PhD-level topics.
Excellent written English and ability to communicate structured, logical explanations.
Proven ability to design challenging problems that probe model reasoning.
Python proficiency for scientific computing and numerical methods.
Familiarity with NumPy, SciPy, SymPy, Pandas, and Matplotlib.
Comfort with numerical methods, algorithms, and remote collaboration.
Helpful background
You don’t need a specific job title to apply, but graduate study (Master’s, PhD, or postdoctoral experience) in physics, applied physics, or a closely related field is strongly helpful. Experience solving and explaining complex physics problems, and a track record of writing structured solutions or teaching, will make you a strong fit.
Experience creating visuals or stepwise derivations to explain physics clearly.
Familiarity with symbolic computation (SymPy) and reproducible numerical verification.
Working details & how it works
This is a contractor, part-time role requiring roughly 20+ hours per week. Work is worldwide and remote; applications in English only. Tasks center on text generation and evaluation rating: you will author test prompts and authoritative solutions, then rate or annotate model outputs against those solutions.
OpenTrain handles the engagement and provides the project scope, evaluation rubrics, and any collaboration with LLM researchers. You will be expected to follow specified guidelines and maintain reproducible, well-documented solutions.
Data type: text; label types: text generation and evaluation/rating.
Employment: contractor, part-time; flexible scheduling but consistent weekly hours required.
Languages: English required; worldwide applicants welcome.
Who should apply and next steps
Apply if you enjoy crafting challenging physics problems, writing clear stepwise solutions, and using Python to verify results. This role suits researchers, instructors, or advanced students who want flexible, impactful work contributing directly to how LLMs learn physics reasoning.
To apply, prepare examples of problems and step-by-step solutions you’ve written (or brief descriptions of similar work). OpenTrain will provide instructions for the next steps in the evaluation and onboarding process.
Join OpenTrain to design and solve advanced physics problems that probe LLM reasoning and symbolic skills; remote, part-time contractor work (20+ hrs/week) for candidates with graduate-level physics experience and strong written English.
Join OpenTrain to design and solve advanced physics problems and build evaluation benchmarks that fine-tune large language models. This remote, part-time contractor role (20+ hrs/week) suits PhD-level physicists or equivalent with strong symbolic and multi-step reasoning skills.
Design and solve advanced physics problems to probe large language models' multi-step reasoning and symbolic skills; remote, contract, 20+ hours/week working with OpenTrain AI to build benchmarks from undergraduate to PhD levels.