Create challenging mathematics problems and rigorous solutions that reveal how large language models handle abstraction, symbolic manipulation, and multi-step reasoning. This remote US freelance role offers 30- or 40-hour weekly commitments.
Generative AI & RLHF
Remote
1 country
Eligibility
Entry
Experience
Aug 30, 2026
Posted
Open to applicants in
United States
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover specialized projects, build a lasting professional profile, and apply to opportunities in minutes. Creating an OpenTrain account is free.
Build a portfolio of AI training and evaluation experience
Find opportunities aligned with your subject-matter expertise
Work remotely on projects shaping how modern AI systems perform
About AI Training and LLM Evaluation
Large language models learn and improve through human-created examples, detailed evaluations, and expert feedback. In this fast-growing field, specialists test model outputs, identify weaknesses, and provide the reasoning that helps AI become more accurate, useful, and reliable.
Work at the intersection of mathematics and cutting-edge AI
Contribute to evaluation benchmarks used to measure model capabilities
Use human judgment to assess reasoning, explanations, and solutions
The Role
OpenTrain is recruiting a Mathematics LLM Evaluation Expert to create and solve challenging mathematical problems that test the capabilities and limitations of large language models. The work covers abstraction, multi-step reasoning, symbolic manipulation, and mathematics topics ranging from early undergraduate study through PhD-level curricula.
This is a remote freelance contractor engagement for specialists based in the United States. The opportunity is listed as entry level, but it requires advanced mathematics knowledge appropriate to graduate or PhD-level study.
Remote freelance contractor role
Candidates must be based in the USA
Available commitment options are 30 or 40 hours per week
The listing indicates a commitment of 20+ hours per week
Continuation may be possible based on performance and project needs
What You'll Do
You will create rigorous evaluation materials and provide detailed feedback to help assess and improve large language models. Your work will combine advanced mathematical problem solving with clear written explanations and careful analysis of reasoning quality.
Design challenging mathematics problems that expose gaps in model reasoning
Explain complex mathematical concepts with accessible language, visuals, and examples
Identify weaknesses in abstraction, multi-step reasoning, and symbolic manipulation
Work with LLM researchers to align problem types with evaluation goals
Contribute to benchmarks based on a broad mathematics curriculum
Provide constructive feedback and detailed annotations that support model improvement
Requirements and Helpful Background
You should have a strong foundation in advanced mathematics and be able to analyze complex problems using a structured, logical approach. Clear, precise communication is essential because solutions, explanations, feedback, and annotations must be understandable and rigorous.
Candidates pursuing or holding a Master's, Ph.D., or postdoctoral degree in Mathematics, Applied Mathematics, Statistics, or a related field are encouraged to apply. Experience developing mathematical explanations, evaluating reasoning quality, or providing detailed academic feedback can support success in this work.
Advanced mathematics knowledge appropriate to graduate or PhD-level study
Strong ability to solve complex problems through structured, logical reasoning
Ability to explain mathematical concepts using simple language, visuals, and examples
Strong English comprehension and structured written communication
Research ability, analytical thinking, and creative problem-solving skills
Ability to provide constructive feedback and detailed annotations
Ability to work independently and collaborate remotely
Reliable computer and internet connection
Working Arrangement
This role is remote and designed for US-based specialists working as freelance contractors. You must be available for at least four hours per day and four hours of overlap with Pacific Time.
Location: United States
Engagement: Freelance contractor
Schedule options: 30 or 40 hours per week
Daily availability: At least four hours
Time-zone overlap: At least four hours with Pacific Time
Language: English
Build Your AI Training Career
Mathematics experts are helping shape how AI systems reason, explain solutions, and handle difficult problems. Through OpenTrain, you can turn this work into credible experience, strengthen your professional profile, and grow a portfolio in AI training and data evaluation.
Apply through OpenTrain in minutes
Showcase specialized mathematics and evaluation experience
Work remotely in a rapidly growing technology field
Create challenging mathematics problems, evaluate large language model reasoning, and build Python and Lean solutions in a flexible remote contractor role. Apply your advanced math expertise to cutting-edge AI training through OpenTrain.
Help advance large language models by designing challenging physics problems, writing rigorous solutions, and shaping evaluation benchmarks. This expert-level remote contract offers 20+ hours per week for graduate-level STEM specialists.
Use advanced physics knowledge to design, solve, and evaluate challenging problems for large language models. This remote, part-time contract role offers 20+ hours per week and welcomes candidates worldwide.