Skip to content
OpenTrain AIFor AI Companies

Math Reasoning Evaluator — Advanced Math Assessment

Evaluate and rate AI-generated mathematical solutions, author exemplar step-by-step answers, and catch subtle conceptual or computational errors; $80/hr, contractor part-time role requiring a BS/MS/PhD in mathematics from a top-100 university and 17–20 hrs/week minimum.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $80/hr

$80/hr

Compensation

Worldwide

Eligibility

Entry

Experience

Oct 24, 2025

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We connect talented contributors with rigorous, paid projects that shape how modern AI systems learn and reason. For this role, OpenTrain AI is the hiring and contracting organization.

Why AI training matters

AI training (also called data labeling or human evaluation) is the human side of building intelligent systems. Experts review model outputs, correct mistakes, and create high-quality examples that teach AI how to reason, explain, and solve problems. This work is remote, flexible, and directly influences state-of-the-art models.

  • Work 100% remotely and shape how AI handles advanced mathematics.
  • Flexible, part-time work that can fit alongside research or other commitments.

The role

You will evaluate AI-generated math responses for correctness, rigor, and clarity. This is a specialist role for mathematically trained candidates (not generalists). Expect to judge proofs, derivations, calculations, and quantitative claims; write model exemplar solutions; and apply detailed rubrics to rate and compare outputs.

  • Position type: Contractor, part-time.
  • Pay: $80 USD per hour.
  • Minimum availability: 17–20 hours per week; preferred cadence ~8 hours/day during active sprints.
  • Work is worldwide and fully remote.

What you'll do day-to-day

Your primary tasks involve careful, expert review of text-based model outputs. Work items usually include multiple short to medium-length math problems and longer proof-style responses.

  • Judge correctness, reasoning depth, and clarity of AI responses using rubrics.
  • Identify and describe subtle conceptual, methodological, or computational errors.
  • Author clear, step-by-step exemplar solutions and model answers.
  • Fact-check quantitative claims and cite reputable public sources when needed.
  • Compare multiple model responses and assign evaluation ratings with consistent justification.

Requirements (must-have)

This role requires proven mathematical expertise and the ability to write rigorous, accessible solutions in strong English (C1+). All must-have qualifications below are required.

  • BS, MS, or PhD in Mathematics, Mathematical Statistics, or Applied Math (completed or in-progress) from a top-100 university.
  • Mastery across core areas (examples: algebra, calculus, probability, statistics) and comfort with formal proofs and notation.
  • Exceptional mathematical writing: lucid, rigorous, stepwise explanations and error analysis.
  • Ability to spot subtle conceptual and computational mistakes and explain them precisely.
  • High attention to detail and consistent application of grading/rating rubrics.
  • Availability for a minimum of 17–20 hours per week.

Preferred experience and bonuses

The following are preferred but not strictly required. They help you be more effective evaluating model reasoning and producing high-quality exemplar solutions.

  • Research experience, analytical writing, or competitive debate experience.
  • Programming literacy (e.g., Python) and facility with LaTeX for clear mathematical typesetting.
  • Prior data labeling, RLHF, or AI model evaluation experience (bonus).

Onboarding, assessments, and schedule

Onboarding includes paid qualification steps to ensure fit and consistency. These are brief, paid exams you must pass before starting active project work.

  • Paid 1–2 hour qualification exam followed by a paid 1–2 hour project exam.
  • Work is organized in sprints; during active sprints a preferred cadence is about 8 hours/day, otherwise minimum weekly commitment remains 17–20 hours.
  • Tasks are text-based evaluations; you will use rubrics and submit written justifications for ratings.

How to apply and who should apply

Apply if you are a mathematically specialized candidate who can write rigorous, stepwise solutions in clear English and commit to the stated weekly hours. This role is aimed at domain specialists rather than general STEM applicants.

  • Ideal applicants: current or former graduate students, researchers, or instructors from strong mathematics programs.
  • You must meet the top-100 university requirement and be comfortable producing precise written feedback and exemplar solutions.
  • If you meet the requirements, expect to take the paid qualification exam as the next step after application.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Mathematics AI Response Evaluation Specialist

Join OpenTrain AI to evaluate and improve AI-generated mathematical answers — $70/hr, 20+ hours/week, remote (selected countries). Ideal for MS/PhD-level mathematicians who can spot errors, write clear model solutions, and rate reasoning quality.

Generative AI & RLHF
Text
Remote · Bangladesh, Bhutan, Brazil +14 more
English
Part-time · Flexible
Entry level
Hourly · $70/hr

Posted Jul 9, 2026

Advanced Mathematics LLM Evaluation Expert

Join OpenTrain AI to design, solve, and evaluate challenging mathematics problems that probe LLM reasoning limits; this remote, part-time contractor role expects 20+ hours/week, advanced graduate-level math knowledge, Python ability, and English fluency.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026

Advanced Mathematics LLM Evaluation Expert

Join OpenTrain AI to design and evaluate graduate- and PhD-level mathematics problems that test and improve large language models; work remotely as a contractor for 20+ hours/week building benchmark questions, reviewing model solutions, writing formal Lean proofs, and validating Python computations.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026