Evaluate advanced mathematics problems, model solutions, computational tasks, and formal proofs to improve language models. Work remotely as a contractor for at least 20 hours per week using Python and Lean.
Generative AI & RLHF
100% Remote
Worldwide
Eligibility
Entry
Experience
Jul 20, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover specialized projects, build a professional profile, and apply in minutes for work that develops into a lasting AI training portfolio.
Creating an OpenTrain account is free, and this opportunity is offered as remote contractor work through OpenTrain.
About AI Training and LLM Evaluation
AI training is the human side of building artificial intelligence. People prepare examples, review model responses, and provide expert feedback so modern AI systems can reason more accurately and communicate more clearly.
In this role, your mathematical expertise will contribute to the evaluation of language models across challenging problems, structured solutions, computational tasks, and formal proofs.
The Role
OpenTrain is recruiting a Mathematics LLM Evaluation Expert for advanced mathematics language model evaluation. You will work with multi-step, abstract, and proof-based mathematics problems, reviewing model-generated reasoning and creating high-quality material for evaluation and training.
The assignment is fully remote contractor work with an expected commitment of at least 20 hours per week. Available commitment options are 20, 30, or 40 hours per week, with at least 4 hours per day and 4 hours of overlap with Pacific Time.
Engagement type: Remote contractor and part-time work
Time requirement: 20 or more hours per week
Available schedules: 20, 30, or 40 hours per week
Daily requirement: At least 4 hours per day
Overlap requirement: 4 hours with Pacific Time
Working language: English
Geographic availability: Worldwide
What You'll Do
You will design, solve, review, and verify advanced mathematics content used to assess language model reasoning. Your work will span topics from early undergraduate mathematics through doctoral-level mathematics and will require careful written explanations and precise technical judgment.
Design original, challenging problems covering multi-step, abstract, and proof-based reasoning.
Solve mathematics problems independently and write logically structured solutions with clear justifications.
Review model-generated solutions for mathematical errors, incomplete reasoning, and missing arguments.
Provide precise feedback, annotations, and corrections on model outputs.
Contribute to evaluation benchmarks spanning early undergraduate through doctoral-level mathematics.
Develop and validate Python solutions using approved scientific libraries.
Use Python for problem solving and numerical verification.
Translate mathematical problems and proofs into Lean.
Verify that formal Lean proofs compile correctly.
Requirements and Preferred Skills
This role requires relevant academic or professional expertise in mathematics, applied mathematics, statistics, or a related field. The work calls for advanced mathematical reasoning and the ability to communicate complex ideas with precision.
Experience with Python and scientific libraries is valuable. Familiarity with Lean and formal proof verification is beneficial, while strong written communication, attention to detail, independent work habits, and the ability to follow detailed project guidelines are required.
Strong foundations in algebra, calculus, analysis, geometry, or topology.
Advanced reasoning across multi-step, abstract, and proof-based problems.
Ability to identify errors and missing arguments in model-generated solutions.
Ability to explain complex mathematics using structured language, visuals, and examples.
Python problem-solving, numerical verification, and scientific-library experience.
Lean theorem-prover skills, including translating and compiling formal proofs.
Strong written communication and careful attention to detail.
Ability to work independently and follow detailed project guidelines.
Why Do AI Training Work
AI training and data-labeling work offers a way to contribute directly to how cutting-edge AI systems are built. Remote projects can provide flexible part-time work for people who want to apply specialized expertise while developing experience in a fast-growing technology field.
By evaluating mathematical reasoning, you help make AI systems more reliable on the kinds of complex problems that demand accuracy, logical structure, and verifiable conclusions.
Fully remote work from anywhere with a computer and internet connection.
Flexible part-time engagement with selectable weekly commitment options.
Opportunity to apply advanced mathematical expertise to modern AI systems.
A chance to build a durable portfolio of specialized AI training experience.
Build Your AI Training Career With OpenTrain
OpenTrain brings opportunities for AI training and data-labeling contributors into one place while helping them build a profile they control. Your experience on specialized projects can support a stronger record of work and help you discover future roles aligned with your expertise.
Apply through OpenTrain to begin building a career at the intersection of mathematics and artificial intelligence.
Help advance large language models by designing challenging physics problems, writing rigorous solutions, and shaping evaluation benchmarks. This expert-level remote contract offers 20+ hours per week for graduate-level STEM specialists.
Use advanced physics knowledge to design challenging problems, solve them step by step, and help evaluate how large language models reason. This flexible remote contractor role is open worldwide.
Help improve large language models by creating and evaluating mathematical reasoning content. Apply your skills in algebra, geometry, calculus, logic, and clear written explanation in a flexible remote contract role.