Join OpenTrain AI to design and evaluate graduate- and PhD-level mathematics problems that test and improve large language models; work remotely as a contractor for 20+ hours/week building benchmark questions, reviewing model solutions, writing formal Lean proofs, and validating Python computations.
Generative AI & RLHF
100% Remote
Worldwide
Eligibility
Entry
Experience
Jul 17, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people start and grow durable freelance careers teaching AI by bringing specialized AI training work, projects, and portfolio-building tools together in one place.
OpenTrain AI is the hiring and contracting organization for this role. We offer flexible, remote contract work that directly influences how state-of-the-art AI systems behave and are evaluated.
Why AI training matters
AI training (also called data labeling or human feedback work) is the human side of building AI: people prepare, evaluate, and correct examples that models learn from. This industry is one of the fastest-growing ways to work in tech, and contributors shape model behavior across domains.
These roles are often remote and flexible, accessible without prior industry employment, and let you apply domain expertise — here, advanced mathematics — to help set the standards for reliable, rigorous model reasoning.
The role
You will design and evaluate challenging mathematics problems and solutions to improve and benchmark large language models. Work focuses on multi-step, abstract, and proof-based mathematics from advanced undergraduate through PhD level.
This is a remote contractor role, part-time (20+ hours/week). Work is collaborative but requires strong independent initiative and excellent written mathematical exposition.
Employment type: Contractor, Part-time
Time commitment: 20+ hours per week
Languages: English
Work location: Remote, worldwide
What you'll do
Your day-to-day tasks blend content creation, critical review, and formal verification. Expect to both produce rigorous reference solutions and evaluate model outputs against those references.
Design original, challenging mathematics problems to probe LLM reasoning limits in multi-step, abstract, and proof-based settings.
Solve problems independently and write detailed, logically structured solutions with clear justifications.
Review model-generated solutions, identify mathematical errors or missing arguments, and provide precise feedback, annotations, and corrections.
Contribute to defining new evaluation benchmarks mapped to mathematics curricula from early undergraduate through PhD topics.
Develop and validate Python-based solutions for computational tasks using approved scientific libraries.
Translate mathematical problems and proofs into Lean formal language and verify that formal proofs compile correctly.
Requirements (must-haves)
Candidates must meet the core technical and operational requirements below. All items come from the role description — we do not invent prerequisites beyond what was provided.
Solid foundation in mathematics at the level expected in engineering entrance exams and graduate or PhD-level programs.
Ability to evaluate multi-step, abstract, and proof-based mathematics at a graduate or PhD level.
Ability to break down complex mathematical concepts into simple, clear explanations and to write detailed, logically structured solutions with clear justifications.
Proven skill identifying mathematical errors in model-generated solutions and providing precise, constructive feedback and annotations.
Ability to develop and validate Python code for computational math tasks (using approved scientific libraries).
Experience translating mathematical problems into Lean formal proof language and verifying formal proofs compile.
Strong research and analytical skills, excellent structured written communication, and comfort collaborating remotely.
Desktop or laptop with a reliable internet connection.
Helpful background (preferred)
The following qualifications are useful but not strictly required. They reflect backgrounds that make the role easier and let you contribute more quickly.
Pursuing or holding a Master's, PhD, or Postdoctoral degree in Mathematics, Applied Mathematics, Statistics, or a related field.
Familiarity with Python and scientific libraries (NumPy, SciPy, SymPy, etc.).
Experience with theorem-proving tools such as Lean and familiarity with formal proof workflows.
How the work is evaluated and delivered
You will complete tasks that include writing reference problems and solutions, rating and annotating model outputs, submitting Python solutions, and producing Lean proofs. Quality is judged on mathematical correctness, clarity of exposition, completeness of proofs, and correctness of formal verification.
Labeling work will include evaluation ratings and text-generation assessments. OpenTrain provides the task interface and guidelines; you will follow project-specific rubrics and review protocols.
Application notes and compensation
Apply with examples of advanced math problems you've written or solved (solutions, notebooks, or Lean files if available) and a short summary of relevant experience. Because this is specialized work, examples of formal proofs, complex problem solutions, or validated Python notebooks are especially helpful.
Compensation details are not listed in this description and are set per project within OpenTrain. This is a contract, part-time engagement with flexible scheduling; specific pay and milestones are provided during onboarding.
Join OpenTrain AI to design, solve, and evaluate challenging mathematics problems that probe LLM reasoning limits; this remote, part-time contractor role expects 20+ hours/week, advanced graduate-level math knowledge, Python ability, and English fluency.
Join OpenTrain to design and solve advanced physics problems that probe LLM reasoning and symbolic skills; remote, part-time contractor work (20+ hrs/week) for candidates with graduate-level physics experience and strong written English.
Join OpenTrain to design and solve advanced physics problems and build evaluation benchmarks that fine-tune large language models. This remote, part-time contractor role (20+ hrs/week) suits PhD-level physicists or equivalent with strong symbolic and multi-step reasoning skills.