Skip to content
OpenTrain AIFor AI Companies

Mathematics LLM Evaluation Expert

Create challenging mathematics problems and rigorous solutions that reveal how large language models handle abstraction, symbolic manipulation, and multi-step reasoning. This remote US freelance role offers 30- or 40-hour weekly commitments.

OpenTrain AI

Generative AI & RLHF

Remote

1 country

Eligibility

Entry

Experience

Aug 30, 2026

Posted

Open to applicants in

United States

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover specialized projects, build a lasting professional profile, and apply to opportunities in minutes. Creating an OpenTrain account is free.

  • Build a portfolio of AI training and evaluation experience
  • Find opportunities aligned with your subject-matter expertise
  • Work remotely on projects shaping how modern AI systems perform

About AI Training and LLM Evaluation

Large language models learn and improve through human-created examples, detailed evaluations, and expert feedback. In this fast-growing field, specialists test model outputs, identify weaknesses, and provide the reasoning that helps AI become more accurate, useful, and reliable.

  • Work at the intersection of mathematics and cutting-edge AI
  • Contribute to evaluation benchmarks used to measure model capabilities
  • Use human judgment to assess reasoning, explanations, and solutions

The Role

OpenTrain is recruiting a Mathematics LLM Evaluation Expert to create and solve challenging mathematical problems that test the capabilities and limitations of large language models. The work covers abstraction, multi-step reasoning, symbolic manipulation, and mathematics topics ranging from early undergraduate study through PhD-level curricula.

This is a remote freelance contractor engagement for specialists based in the United States. The opportunity is listed as entry level, but it requires advanced mathematics knowledge appropriate to graduate or PhD-level study.

  • Remote freelance contractor role
  • Candidates must be based in the USA
  • Available commitment options are 30 or 40 hours per week
  • The listing indicates a commitment of 20+ hours per week
  • Continuation may be possible based on performance and project needs

What You'll Do

You will create rigorous evaluation materials and provide detailed feedback to help assess and improve large language models. Your work will combine advanced mathematical problem solving with clear written explanations and careful analysis of reasoning quality.

  • Design challenging mathematics problems that expose gaps in model reasoning
  • Develop accurate, detailed, step-by-step solutions
  • Explain complex mathematical concepts with accessible language, visuals, and examples
  • Identify weaknesses in abstraction, multi-step reasoning, and symbolic manipulation
  • Work with LLM researchers to align problem types with evaluation goals
  • Contribute to benchmarks based on a broad mathematics curriculum
  • Provide constructive feedback and detailed annotations that support model improvement

Requirements and Helpful Background

You should have a strong foundation in advanced mathematics and be able to analyze complex problems using a structured, logical approach. Clear, precise communication is essential because solutions, explanations, feedback, and annotations must be understandable and rigorous.

Candidates pursuing or holding a Master's, Ph.D., or postdoctoral degree in Mathematics, Applied Mathematics, Statistics, or a related field are encouraged to apply. Experience developing mathematical explanations, evaluating reasoning quality, or providing detailed academic feedback can support success in this work.

  • Advanced mathematics knowledge appropriate to graduate or PhD-level study
  • Strong ability to solve complex problems through structured, logical reasoning
  • Ability to explain mathematical concepts using simple language, visuals, and examples
  • Strong English comprehension and structured written communication
  • Research ability, analytical thinking, and creative problem-solving skills
  • Ability to provide constructive feedback and detailed annotations
  • Ability to work independently and collaborate remotely
  • Reliable computer and internet connection

Working Arrangement

This role is remote and designed for US-based specialists working as freelance contractors. You must be available for at least four hours per day and four hours of overlap with Pacific Time.

  • Location: United States
  • Engagement: Freelance contractor
  • Schedule options: 30 or 40 hours per week
  • Daily availability: At least four hours
  • Time-zone overlap: At least four hours with Pacific Time
  • Language: English

Build Your AI Training Career

Mathematics experts are helping shape how AI systems reason, explain solutions, and handle difficult problems. Through OpenTrain, you can turn this work into credible experience, strengthen your professional profile, and grow a portfolio in AI training and data evaluation.

  • Apply through OpenTrain in minutes
  • Showcase specialized mathematics and evaluation experience
  • Work remotely in a rapidly growing technology field

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Advanced Mathematics LLM Evaluation Expert

Create challenging mathematics problems, evaluate large language model reasoning, and build Python and Lean solutions in a flexible remote contractor role. Apply your advanced math expertise to cutting-edge AI training through OpenTrain.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026

Physics LLM Evaluation Expert

Help advance large language models by designing challenging physics problems, writing rigorous solutions, and shaping evaluation benchmarks. This expert-level remote contract offers 20+ hours per week for graduate-level STEM specialists.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Expert level

Posted Jul 17, 2026

Physics Problem Designer for LLM Evaluation

Use advanced physics knowledge to design, solve, and evaluate challenging problems for large language models. This remote, part-time contract role offers 20+ hours per week and welcomes candidates worldwide.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026