Create advanced physics problems, solve them with clear step-by-step reasoning, and evaluate large language model outputs. This remote 12-week contract requires at least 20 hours weekly and four hours of Pacific Time overlap.
The work
You will help evaluate large language models by creating difficult physics problems and writing accurate, detailed solutions. Your work will support benchmark development from early undergraduate topics through PhD-level physics.
You will also review model responses, identify weaknesses in complex reasoning, and provide annotations that help improve evaluation quality.
- Design challenging physics problems that test advanced reasoning.
- Solve problems accurately and write clear, step-by-step solutions for gold-standard data.
- Create new scenarios that reveal weaknesses in multi-step reasoning.
- Work with LLM researchers to align tasks with evaluation goals.
- Help build physics evaluation benchmarks across several curriculum levels.
- Give constructive feedback and detailed annotations on model outputs.
What it pays and takes
This is a remote contractor engagement for 12 weeks. The role is open worldwide, and you must be comfortable working in English.
The posting does not provide a pay rate. Prior AI evaluation or annotation experience is useful but not required.
- Engagement: 12-week remote contract.
- Time: At least 4 hours per day and up to 40 hours per week; the listed commitment is 20+ hours weekly.
- Schedule: At least 4 hours of overlap with Pacific Time.
- Education: A background or doctorate in mathematics, physics, or an equivalent technical field.
- Skills: Strong analytical, research, and problem-solving ability.
- Communication: Excellent English comprehension and structured written communication.
- Reasoning: Ability to explain complex mathematical or physical ideas clearly and step by step.
- Work setup: Ability to work independently with a reliable computer and high-speed internet.
- Helpful experience: AI evaluation, data annotation, content review, quality assurance, or a related analytical role.
How it works
Apply on OpenTrain with your resume, then complete the application on the hiring site.
About AI training work
AI training is the human work behind systems that generate and understand text, images, code, and other content. People create examples, review model responses, and explain what a strong answer should look like, so deep subject knowledge is valuable.
OpenTrain is where people find and build careers in AI training and data labeling. Creating an account is free.