Skip to content
OpenTrain AIFor AI Companies

AI Evaluation Engineer, Engineering Simulation

Create and validate challenging engineering simulation benchmarks that train and evaluate AI agents. This remote contractor role combines advanced engineering design, Python, open-source simulation, and model failure analysis.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Entry

Experience

Sep 11, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. OpenTrain helps specialists discover cutting-edge projects, build a durable professional profile, and apply for opportunities in minutes.

Creating an OpenTrain account is free. Your profile can help showcase credible AI training and evaluation experience as you grow.

About AI Training and Model Evaluation

AI training is the human side of building artificial intelligence. Engineers, researchers, and other specialists create examples, review model behavior, and evaluate outputs so modern AI systems can reason more reliably.

In this role, your engineering judgment will help assess AI agents working through realistic, simulation-based design problems. The work is part of a fast-growing field where human expertise directly shapes how advanced systems perform.

The Role

OpenTrain is recruiting an AI Evaluation Engineer focused on engineering simulation and design. As a contractor, you will create and validate demanding engineering problems used to train and evaluate AI agents.

The work spans electrical, mechanical, control systems, aerospace, systems, and robotics applications. You will combine rigorous engineering judgment with benchmark development, simulation, coding-agent evaluation, and failure analysis.

  • Remote contractor engagement
  • Expected duration of up to 24 weeks
  • 40 hours per week, with four hours of overlap with PST
  • Part-time engagement may be acceptable
  • Weekend on-call availability required
  • Stable high-speed internet and a suitable computer required
  • Compensation is not stated in the role

What You’ll Do

You will develop technically rigorous evaluation environments and inspect how AI coding agents approach complex engineering tasks. The goal is to produce benchmarks that are challenging, clear, physically plausible, and objectively measurable.

  • Author self-contained engineering design tasks with competing constraints, explicit optimization goals, validated reference solutions, and objective automated graders.
  • Build, run, and validate simulation environments with open-source tools and custom Python test benches.
  • Review coding-agent outputs and execution logs across repeated trials to identify systemic reasoning, tool-use, and simulator-feedback failures.
  • Refine benchmark difficulty using empirical model performance data while preserving clarity and technical validity.
  • Collaborate with AI researchers and engineering specialists to integrate high-rigor benchmarks into evaluation workflows.
  • Assess design judgment across electrical, mechanical, control systems, aerospace, systems, or robotics domains.
  • Evaluate physically plausible problems with valid boundary conditions, consistent units, convergence criteria, and optimization targets.
  • Analyze pass@k results, trajectory logs, nondeterministic behavior, and systemic failure modes.

Required Qualifications

This opportunity is listed as entry level, but the role requires substantial specialist preparation and hands-on engineering experience. Candidates should be comfortable combining advanced engineering design with software-based simulation and AI evaluation.

  • Master’s degree or PhD in electrical, mechanical, aerospace, control systems, systems engineering, robotics, or an applied science discipline.
  • Three or more years of hands-on engineering design experience.
  • Strong Python scripting ability.
  • Proficiency with at least one relevant open-source simulation package, such as ngspice, PySpice, OpenFOAM, FEniCSx, CalculiX, python-control, CadQuery, build123d, OpenModelica, Cantera, or Gmsh.
  • Practical experience with LLMs or coding agents, pass@k, failure-mode analysis, nondeterministic behavior, and trajectory-log review.
  • Careful reasoning about physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation.

Who Should Apply

This role may suit an engineer who enjoys turning real design constraints into precise computational problems and investigating why AI systems succeed or fail. It is especially relevant for specialists with experience in simulation, engineering design, Python automation, or robotics and control workflows.

  • Engineering specialists who can assess advanced design decisions across multiple technical domains.
  • Python developers with practical experience building test benches or simulation workflows.
  • Researchers and practitioners familiar with LLM evaluation, coding agents, or benchmark design.
  • Candidates who can document assumptions and verify units, constraints, physical behavior, and convergence.

Engagement Details

The engagement is worldwide and remote, with English as the working language. The expected commitment is 20 or more hours per week in the role information, with the schedule specifying 40 hours per week and allowing part-time engagement in some cases.

Applicants should be prepared for four hours of overlap with PST and weekend on-call availability. A stable high-speed internet connection and suitable computer are required.

  • Work location: Worldwide and remote
  • Language: English
  • Engagement type: Contractor and part-time
  • Time requirement: 20 or more hours per week; standard schedule is 40 hours per week
  • PST overlap: Four hours required
  • Weekend on-call availability: Required
  • Duration: Up to 24 weeks

How to Apply Through OpenTrain

Create a free OpenTrain account, build your profile around your engineering and AI evaluation experience, and apply in minutes. OpenTrain is designed to help contributors find and grow in AI training and data-labeling work while maintaining a credible portfolio of specialized experience.

  • Highlight your engineering design and simulation background.
  • List relevant Python and open-source simulation tools.
  • Describe experience with LLMs, coding agents, pass@k, or trajectory-log analysis.
  • Show how you validate physical plausibility, constraints, optimization goals, and technical documentation.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Engineering Simulation AI Evaluation Expert

Use deep engineering expertise to build Python simulations, define quantitative acceptance criteria, and evaluate whether AI-generated designs work from first principles. Remote contract work pays $60 to $90 per hour.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $60–$90/hr

Posted Sep 1, 2026

Mechanical Engineering AI Training Expert

Use mechanical engineering, simulation, and Python expertise to create, solve, and validate tasks for AI training. Work remotely for about 15 hours per week at $80 to $130 per hour.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $80–$130/hr

Posted Sep 7, 2026

MCP AI Software Evaluation Engineer

Evaluate AI agents on realistic software engineering tasks and build reliable MCP-based reinforcement learning environments. This flexible remote contractor role offers an ideal rate range of $60 to $120 per hour.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $60–$120/hr

Posted Aug 25, 2026