Skip to content
OpenTrain AIFor AI Companies

Staff Research Engineer, Frontier AI Evaluation

Work 20+ hours per week on frontier AI research, evaluation, agent environments, synthetic data, and post-training. This worldwide contractor role requires advanced research engineering experience and English fluency.

Apply now
OpenTrain AI

Generative AI & RLHF

100% Remote

Worldwide

Eligibility

Expert

Experience

Sep 30, 2026

Posted

Open worldwide

The Work

As a Staff Research Engineer, you will investigate the capabilities, limits, and training methods of advanced AI systems. You will combine research and engineering to turn reliable findings into practical AI applications.

You will work across research-grade datasets, reinforcement learning environments, model evaluation, synthetic and agentic data generation, benchmarks, and model understanding. The work may include coding-agent, computer-use, browser-use, and function-calling systems.

  • Formulate research questions and design experiments that produce evidence-based conclusions.
  • Build datasets, prototypes, tools, benchmarks, and evaluation frameworks for complex, multi-step workflows.
  • Train, test, and evaluate models with modern machine learning tools, then analyze the results.
  • Explore reinforcement learning, post-training, synthetic data, agentic systems, model understanding, and AI evaluation.
  • Collaborate with research, engineering, product, and operations stakeholders.
  • Share findings through technical reports or other appropriate channels, and contribute to peer review, technical discussions, mentoring, and research collaboration.

What It Pays And Takes

The listing does not state a pay rate. This is a part-time contractor role for an expert-level candidate, with a commitment of at least 20 hours per week.

  • Pay: Not provided in the listing.
  • Workload: 20+ hours per week.
  • Location: Worldwide.
  • Language: English.
  • Education: A PhD or master's degree in artificial intelligence, machine learning, computer science, or a closely related technical field, or equivalent research experience.
  • Experience: At least 7 years of professional experience, including substantial research engineering work in machine learning or frontier AI systems.
  • Technical skills: Strong Python programming and experience with modern artificial intelligence and machine learning frameworks.
  • Research skills: Experience designing experiments, training or evaluating models, or developing AI systems.
  • Working style: Sound judgment about experimental rigor and reproducibility, clear communication, independent execution, cross-functional collaboration, and technical mentoring experience.
  • Valuable background: Synthetic or agentic data generation, reinforcement learning, post-training, model understanding, AI evaluation, benchmarks, agents, tool-using systems, coding-agent environments, user-interface or browser-use environments, function-calling systems, technical publications, open-

How It Works

Apply on OpenTrain with your resume, then complete the application on the hiring site.

About AI Training Work

AI training is the human work behind modern AI systems, including preparing datasets, testing models, rating outputs, and building evaluation methods. Experienced researchers are paid to design reliable experiments, improve data and evaluation quality, and help determine how advanced systems perform in real tasks.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Market Research AI Evaluation Expert

Review AI-generated market research and create expert research deliverables, including surveys, competitor scans, and insight summaries. Earn $30-$65 per hour with 20+ hours per week in a remote contractor role.

Generative AI & RLHF
Text
Remote · Bangladesh, Bhutan, Brazil +14 more
English
Part-time · Flexible
Intermediate level
Hourly · $30–$65/hr

Posted Jul 9, 2026

Personalized AI Response Evaluation Analyst

Create personal-context prompts, compare AI responses, and explain issues such as unsupported claims or forced connections. This remote, entry-level contract role requires 20+ hours weekly and strong English writing.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Sep 4, 2026

AI Evaluation Question Development Specialist

Create difficult question-and-answer pairs that test advanced AI models, supported by citations and clear reasoning. This remote contractor role pays $40 to $90 per hour and requires strong research experience and written English.

Generative AI & RLHF
Text
Remote · Andorra, United Arab Emirates, Antigua & Barbuda +227 more
English
Part-time · Flexible
Entry level
Hourly · $40–$90/hr

Posted Sep 8, 2026