Math Reasoning Evaluator — Advanced Math Assessment
Evaluate and rate AI-generated mathematical solutions, author exemplar step-by-step answers, and catch subtle conceptual or computational errors; $80/hr, contractor part-time role requiring a BS/MS/PhD in mathematics from a top-100 university and 17–20 hrs/week minimum.
Generative AI & RLHF
100% Remote Hourly · $80/hr
$80/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Oct 24, 2025
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We connect talented contributors with rigorous, paid projects that shape how modern AI systems learn and reason. For this role, OpenTrain AI is the hiring and contracting organization.
Why AI training matters
AI training (also called data labeling or human evaluation) is the human side of building intelligent systems. Experts review model outputs, correct mistakes, and create high-quality examples that teach AI how to reason, explain, and solve problems. This work is remote, flexible, and directly influences state-of-the-art models.
Work 100% remotely and shape how AI handles advanced mathematics.
Flexible, part-time work that can fit alongside research or other commitments.
The role
You will evaluate AI-generated math responses for correctness, rigor, and clarity. This is a specialist role for mathematically trained candidates (not generalists). Expect to judge proofs, derivations, calculations, and quantitative claims; write model exemplar solutions; and apply detailed rubrics to rate and compare outputs.
Position type: Contractor, part-time.
Pay: $80 USD per hour.
Minimum availability: 17–20 hours per week; preferred cadence ~8 hours/day during active sprints.
Work is worldwide and fully remote.
What you'll do day-to-day
Your primary tasks involve careful, expert review of text-based model outputs. Work items usually include multiple short to medium-length math problems and longer proof-style responses.
Judge correctness, reasoning depth, and clarity of AI responses using rubrics.
Identify and describe subtle conceptual, methodological, or computational errors.
Author clear, step-by-step exemplar solutions and model answers.
Fact-check quantitative claims and cite reputable public sources when needed.
Compare multiple model responses and assign evaluation ratings with consistent justification.
Requirements (must-have)
This role requires proven mathematical expertise and the ability to write rigorous, accessible solutions in strong English (C1+). All must-have qualifications below are required.
BS, MS, or PhD in Mathematics, Mathematical Statistics, or Applied Math (completed or in-progress) from a top-100 university.
Mastery across core areas (examples: algebra, calculus, probability, statistics) and comfort with formal proofs and notation.
Exceptional mathematical writing: lucid, rigorous, stepwise explanations and error analysis.
Ability to spot subtle conceptual and computational mistakes and explain them precisely.
High attention to detail and consistent application of grading/rating rubrics.
Availability for a minimum of 17–20 hours per week.
Preferred experience and bonuses
The following are preferred but not strictly required. They help you be more effective evaluating model reasoning and producing high-quality exemplar solutions.
Research experience, analytical writing, or competitive debate experience.
Programming literacy (e.g., Python) and facility with LaTeX for clear mathematical typesetting.
Prior data labeling, RLHF, or AI model evaluation experience (bonus).
Onboarding, assessments, and schedule
Onboarding includes paid qualification steps to ensure fit and consistency. These are brief, paid exams you must pass before starting active project work.
Paid 1–2 hour qualification exam followed by a paid 1–2 hour project exam.
Work is organized in sprints; during active sprints a preferred cadence is about 8 hours/day, otherwise minimum weekly commitment remains 17–20 hours.
Tasks are text-based evaluations; you will use rubrics and submit written justifications for ratings.
How to apply and who should apply
Apply if you are a mathematically specialized candidate who can write rigorous, stepwise solutions in clear English and commit to the stated weekly hours. This role is aimed at domain specialists rather than general STEM applicants.
Ideal applicants: current or former graduate students, researchers, or instructors from strong mathematics programs.
You must meet the top-100 university requirement and be comfortable producing precise written feedback and exemplar solutions.
If you meet the requirements, expect to take the paid qualification exam as the next step after application.
Join OpenTrain AI to evaluate and improve AI-generated mathematical answers — $70/hr, 20+ hours/week, remote (selected countries). Ideal for MS/PhD-level mathematicians who can spot errors, write clear model solutions, and rate reasoning quality.
Join OpenTrain AI to design, solve, and evaluate challenging mathematics problems that probe LLM reasoning limits; this remote, part-time contractor role expects 20+ hours/week, advanced graduate-level math knowledge, Python ability, and English fluency.
Join OpenTrain AI to design and evaluate graduate- and PhD-level mathematics problems that test and improve large language models; work remotely as a contractor for 20+ hours/week building benchmark questions, reviewing model solutions, writing formal Lean proofs, and validating Python computations.