Create rigorous, executable AI evaluation tasks grounded in scientific computing, mathematics, and programming. This remote contract role pays $70 per hour and requires 20+ hours weekly.
Coding & Software
100% Remote Hourly · $70/hr
$70/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Aug 27, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring contractors to help shape how advanced AI systems solve challenging technical problems.
Creating an OpenTrain account is free, and contributors can build a profile that showcases their AI training experience and technical strengths.
Remote contract opportunity available worldwide
Part-time schedule with a commitment of 20+ hours per week
Pay rate: $70 per hour
Work language: English
About AI Evaluation Work
AI evaluation is the human side of improving modern artificial intelligence. Specialists create examples, assess model outputs, and define what a correct, useful, and technically rigorous answer looks like.
In this role, your scientific judgment will help test whether frontier models can produce reliable solutions to executable research problems.
Contribute to cutting-edge AI development
Apply advanced technical knowledge to real evaluation tasks
Help establish standards for correctness and scientific quality
The Role
OpenTrain is seeking a Scientific Computing AI Evaluation Task Author to develop challenging evaluation tasks focused on executable research problems. You will transform authentic mathematical and computational material into prompts and grading standards that test whether AI models can produce correct solutions.
The work combines scientific judgment, programming, and rigorous evaluation across numerical linear algebra, computational mechanics, and computational finance.
Contractor position
Part-time engagement
Entry-level classification with PhD-level technical requirements
20+ hours per week
$70 per hour
What You'll Do
You will source suitable material, design original scientific challenges, and evaluate model performance. Tasks should have a clear computational focus and be sufficiently demanding to distinguish correct technical reasoning from plausible but incorrect answers.
Authoring work will be maintained through pull requests and automated quality checks, so careful documentation and dependable technical workflows are important.
Source material from published papers, Kaggle datasets, open-source repositories, or original scenarios
Write original scientific prompts with a clear computational focus
Create grading criteria that define what constitutes a correct answer
Calibrate tasks against frontier models to keep released evaluations genuinely challenging
Maintain task-authoring work through pull requests and automated quality checks
Required Qualifications
This role requires advanced academic training and the ability to judge the correctness of technical answers. Candidates must have a PhD in mathematics, applied mathematics, computational mathematics, or a closely related field.
You must also demonstrate depth in at least two of numerical linear algebra, computational mechanics, and computational finance.
PhD-level training in mathematics, applied mathematics, computational mathematics, or a closely related field
Depth in at least two of numerical linear algebra, computational mechanics, and computational finance
Ability to write executable scientific problems and rigorous grading criteria
Working proficiency in Python or R for scientific computing
Practical experience with GitHub and pull-request workflows
Working knowledge of Docker
Ability to evaluate the correctness of technical answers
Helpful Background
The following experience is helpful but not listed as required. It may be especially useful when sourcing authentic research material or designing executable problems that reflect professional scientific practice.
Peer-reviewed journal publications
Scientific software development experience
Research engineering experience
Why Work in AI Training
AI training and data labeling work supports the development of modern AI systems through carefully prepared examples and expert human feedback. Contributors with specialized knowledge can work directly on how state-of-the-art models reason, code, and respond.
Many projects are remote and flexible, making this work compatible with other professional or academic commitments while offering a way to build experience in a rapidly growing technology field.
Work remotely from anywhere with an internet connection
Choose flexible work that can fit around other commitments
Use specialized expertise to influence advanced AI systems
Build experience in a fast-growing area of technology
How to Apply
Apply through OpenTrain to be considered for this Scientific Computing AI Evaluation Task Author contract. Your profile can help present your academic background, scientific computing capabilities, and relevant research or software experience in one place.
Create a free OpenTrain account
Complete your profile with relevant qualifications and experience
Apply in minutes through OpenTrain
Be prepared to demonstrate your scientific computing and evaluation abilities
Use advanced physics and scientific programming to create executable research problems that test frontier AI models. This remote six-week contract pays $70 per hour for at least 20 hours weekly.
Create challenging AI evaluation tasks focused on semiconductor materials and molecular modeling. Use scientific judgment, programming, and rubric design in a flexible, worldwide contract role paying $70 per hour.
Use your PhD-level chemistry expertise to create and evaluate challenging scientific computing tasks for frontier AI models. This remote, six-week contract pays $70 per hour for 20 or more hours weekly.