Computational Structural Engineering AI Benchmark Designer
Create graduate-level computational engineering benchmarks that test advanced AI models on real scientific software workflows. Use Python, Linux, and engineering expertise in a flexible worldwide contractor role paying $70 to $85 per hour.
Coding & Software
100% Remote Hourly · $70–$85/hr
$70–$85/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Aug 21, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring contractors to create specialized evaluation work that helps shape how advanced AI systems reason about real technical problems.
This opportunity is available worldwide as a part-time contractor engagement. Creating an OpenTrain account is free, and contributors can use the platform to build a profile and apply to AI training projects.
Worldwide opportunity
Contractor and part-time engagement
Apply through OpenTrain
About AI Benchmark Design
AI training is the human side of building artificial intelligence. In this role, you will design and evaluate challenging computational tasks so AI models can be tested on authentic scientific software workflows rather than simple text exercises.
Your work will help measure whether models can execute complex multi-step calculations, plan efficient queries or experiments, and produce reliable technical results.
Work at the intersection of engineering, software, and AI evaluation
Create research-level tasks based on real computational workflows
Contribute to the development of more capable AI systems
The Role
OpenTrain is seeking a Computational Structural and Mechanical Engineering Expert to design original benchmark problems for advanced AI models. Problems may require exact answers from fully defined setups or strategic planning to uncover hidden information through a sequence of queries or experiments.
Each benchmark is tested against state-of-the-art AI models. You will refine the problem design and difficulty based on test results and feedback, then implement the supporting technical components in Python.
Hourly compensation: $70 to $85 USD
Listed rate: $85 USD per hour
At least 15 to 20 hours per week, with a default commitment of 40 hours per week
English-language work
Entry-level listing with graduate-level specialist requirements
What You'll Do
You will create computational problems that require skilled use of specialized open-source scientific software. The work involves building complete, testable evaluation tasks and adjusting them until they meet the intended challenge level.
You will also write the Python logic needed to establish expected results and assess model submissions, while working through testing loops with remote compute sandboxes.
Create original graduate-level structural and mechanical engineering problems
Develop problems with exact answers from fully defined setups
Design query or experiment planning tasks that use partial results strategically
Write problem setups, oracle functions, and solution validators in Python
Test benchmarks against advanced AI models
Iterate on difficulty and problem design using evaluation feedback
Required Qualifications
You should have graduate-level training in a relevant STEM field, such as an MS, PhD, or equivalent research experience, with expertise in computational structural or mechanical engineering.
Strong programming ability and practical experience with scientific computing are essential. You should also be comfortable working independently in technical environments where problems and evaluation logic must be developed and refined.
Graduate-level expertise in computational structural or mechanical engineering
MS, PhD, or equivalent research experience in a relevant STEM field
Hands-on proficiency with at least one listed open-source computational tool
Strong Python skills for problem setups, oracle functions, and solution validators
Ability to design challenging evaluation problems that test AI model reasoning
Comfort working independently and incorporating feedback
Experience with Linux, terminal tools, and remote compute sandboxes
Supported Scientific Tools
You must have proven hands-on proficiency with at least one open-source computational tool from the list below. Experience with multiple tools or domains is helpful but not required.
FEniCSx or DOLFINx
scikit-fem, deal.II, MFEM, or MOOSE
CalculiX, Elmer FEM, Code_Aster, or SfePy
FiPy, Devito, Cantera, CoolProp, Pyomo, or SimPy
Helpful Background
The strongest candidates may bring experience beyond the core requirements, particularly in benchmark development, scientific education, or reproducible computational research.
Experience across multiple relevant computational domains or tools
Familiarity with benchmark or evaluation design
Scientific teaching or exam and problem-set design experience
Experience with computational reproducibility
Familiarity with containerized environments
How This Work Fits AI Careers
AI training and data labeling work includes evaluating model outputs, writing examples, and testing whether systems can solve demanding technical tasks. Unlike traditional annotation, this project calls for deep engineering knowledge and hands-on software development.
The work is remote-friendly and flexible, making it possible to contribute part time while applying specialized expertise to the fast-growing field of AI development.
Use advanced engineering expertise to influence AI evaluation
Work flexibly as a part-time contractor
Build experience in a growing technical AI training field
Apply through OpenTrain and develop your AI training career
Create demanding particle and nuclear physics benchmarks that test whether advanced AI systems can perform research-level computational work. Use Python, scikit-hep, and scientific workflows in a remote contract role paying $70 to $100 per hour.
Create and solve research-grade computational mathematics problems that help evaluate advanced language models. Use Python, rigorous reasoning, and technical review skills in a flexible 20+ hour weekly contract.
Use advanced computer science expertise to author and review rigorous benchmark questions for AI research, including solutions, distractors, and academic references. This fully remote contract pays $66 to $84 per hour.