Join OpenTrain AI to design, solve, and evaluate challenging mathematics problems that probe LLM reasoning limits; this remote, part-time contractor role expects 20+ hours/week, advanced graduate-level math knowledge, Python ability, and English fluency.
Generative AI & RLHF
100% Remote
Worldwide
Eligibility
Entry
Experience
Jul 20, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people building careers in AI training and data labeling. We connect skilled contributors with specialized, cutting-edge projects so you can build a durable freelance career teaching AI.
OpenTrain AI is the hiring and contracting organization for this role. Creating an OpenTrain account is free and lets you find projects, build a unified portfolio, and apply quickly.
About AI training and why this work matters
AI training is the human side of building intelligent systems: people create, check, and refine examples that modern models learn from. Contributors directly shape how models reason, solve problems, and behave in real-world tasks.
This role focuses on evaluation and benchmark creation for mathematical reasoning—an area that helps improve model accuracy on multi-step, abstract, and proof-oriented problems.
The role
You will be an Advanced Mathematics LLM Evaluation Expert: designing original problems, solving them, reviewing model outputs, and contributing to new evaluation benchmarks. This is a remote contractor position for OpenTrain AI.
Workload: 20+ hours per week. Language: English. Employment types: contractor, part-time. This role is open worldwide.
Employment: Contractor, part-time
Weekly time expectation: 20+ hours/week
Location: Remote, worldwide (English required)
What you'll do day-to-day
Your primary focus is creating and curating high-quality math evaluation content and assessing model outputs against rigorous mathematical standards.
Design original, multi-step, and proof-based mathematics problems that push LLM reasoning limits
Solve problems independently and write clear, logically structured solutions with step-by-step justifications
Review model-generated solutions, identify mathematical errors, and provide precise corrective feedback and annotations
Develop and validate Python-based solutions for computational tasks using approved scientific libraries
Translate problems and proofs into formal language for theorem-prover tasks using Lean and evaluate formal proofs
Requirements
Candidates must demonstrate strong advanced-mathematics ability, excellent written communication, and the ability to work independently in a remote contracting role.
Strong foundation in advanced mathematics at graduate or PhD level (algebra, analysis, topology, geometry, etc.)
Ability to analyze and solve complex, multi-step mathematical problems with a structured approach
Excellent written communication and skill at explaining math clearly in simple language
Proven ability to review and annotate model-generated solutions with precise mathematical feedback
Proficiency using a desktop or laptop with a reliable internet connection
Helpful background (not strictly required)
The following experience will help you contribute immediately and take on more advanced benchmark and theorem-prover tasks.
Master's, PhD, or postdoctoral experience in Mathematics, Applied Mathematics, Statistics, or a related field
Experience writing Python solutions and using scientific libraries for verification and computation
Familiarity with theorem provers such as Lean and translating informal proofs into formal statements
How to apply and what to expect
Create a free OpenTrain account, build a profile that highlights your math background and technical skills, and submit your application for this role. OpenTrain AI handles contracting and project assignments.
We will evaluate applications based on math expertise, clarity of written solutions, and relevant tooling experience (Python, Lean). Specific project or assessment details will be shared after application.
Application step: create an OpenTrain account (free) and submit your profile and materials
Selection factors: demonstrated math ability, clear written solutions, Python or theorem-prover experience
Join OpenTrain AI to design and evaluate graduate- and PhD-level mathematics problems that test and improve large language models; work remotely as a contractor for 20+ hours/week building benchmark questions, reviewing model solutions, writing formal Lean proofs, and validating Python computations.
Join OpenTrain to design and solve advanced physics problems that probe LLM reasoning and symbolic skills; remote, part-time contractor work (20+ hrs/week) for candidates with graduate-level physics experience and strong written English.
Join OpenTrain to design and solve advanced physics problems and build evaluation benchmarks that fine-tune large language models. This remote, part-time contractor role (20+ hrs/week) suits PhD-level physicists or equivalent with strong symbolic and multi-step reasoning skills.