Physics LLM Evaluation Expert — STEM Problem Design
Join OpenTrain to design and solve advanced physics problems and build evaluation benchmarks that fine-tune large language models. This remote, part-time contractor role (20+ hrs/week) suits PhD-level physicists or equivalent with strong symbolic and multi-step reasoning skills.
Generative AI & RLHF
100% Remote
Worldwide
Eligibility
Expert
Experience
Jul 17, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the leading platform for people who build careers in AI training and data labeling. We connect skilled contributors with projects that shape how state-of-the-art models behave and let you consolidate experience, build a visible portfolio, and grow a durable freelance career.
OpenTrain AI is the hiring organization for this role; creating an OpenTrain account is free and required to apply.
Why AI training matters (and why it's a great fit)
AI training is the human work behind modern machine learning: people design tasks, create high-quality answers, and evaluate model outputs so models learn correct, reliable behavior. Contributors often work remotely, set flexible hours, and directly influence cutting-edge systems.
This role places you at the intersection of physics and model evaluation — your technical judgment and carefully written solutions will help fine-tune language models on challenging STEM reasoning.
100% remote work that can fit around other commitments
Flexible, part-time, and contract-friendly opportunities
Work that directly shapes how LLMs reason about advanced STEM topics
Role overview — what you'll do
As a Physics LLM Evaluation Expert you will design hard STEM problems, produce clear step-by-step solutions, and help define benchmarks that probe model limits across physics curricula from early undergraduate through PhD-level material. You will collaborate with LLM researchers to ensure tasks match evaluation goals and produce high-quality annotations and ratings.
Design and solve challenging physics and STEM problems that probe model limitations
Write detailed, step-by-step solutions with clear reasoning and symbolic work
Create and refine new evaluation benchmarks informed by physics curricula
Collaborate with researchers to align tasks, annotation guidelines, and evaluation goals
Provide constructive annotations and evaluation ratings for model outputs (text generation and evaluation)
Requirements
You must be able to work independently in a remote contractor role for 20+ hours per week and communicate clearly in English. The role emphasizes rigorous, structured explanations and high-quality annotation work.
Graduate-level experience in STEM; typically graduate, Ph.D., or postdoctoral background in physics or closely related fields
Strong analytical, research, and multi-step reasoning abilities
Excellent English comprehension and structured written communication
Ability to produce detailed, step-by-step solutions and annotations
Experience with abstraction and symbolic manipulation
Comfort providing constructive feedback and evaluation ratings
Helpful background and strengths
We welcome applicants who have hands-on experience explaining complex physics concepts clearly and concisely, whether through teaching, research, or writing. Visuals and structured reasoning are valuable skills but not strictly required.
Applied physics or related STEM specialization
Prior experience creating curriculum-level problems, exam questions, or research problem sets
Experience turning technical solutions into clear, accessible explanations
Familiarity with benchmark design or model evaluation workflows is a plus
Join OpenTrain to design and solve advanced physics problems that probe LLM reasoning and symbolic skills; remote, part-time contractor work (20+ hrs/week) for candidates with graduate-level physics experience and strong written English.
Design and solve advanced physics problems to probe large language models' multi-step reasoning and symbolic skills; remote, contract, 20+ hours/week working with OpenTrain AI to build benchmarks from undergraduate to PhD levels.
Design and solve challenging physics problems and write step-by-step solutions to probe and evaluate large language models across undergraduate to PhD topics. Remote, part-time contract (~20+ hrs/week); requires strong physics foundations and Python scientific-computing skills.