Skip to content
OpenTrain AIFor AI Companies

AI Evaluation Question Development Specialist

Use advanced research expertise to create challenging, evidence-based questions that evaluate how AI models reason and synthesize information. This remote contractor role pays $40-$90 per hour and requires 20+ hours weekly.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $40–$90/hr

$40–$90/hr

Compensation

Worldwide

Eligibility

Entry

Experience

Sep 8, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this opportunity. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes.

Creating an OpenTrain account is free, and the platform helps AI training professionals develop a lasting portfolio of work in a rapidly growing technology field.

About AI Model Evaluation Work

AI training is the human side of building artificial intelligence. Researchers and specialists create examples, questions, ratings, and feedback that help modern AI systems become more accurate, useful, and capable.

In this role, your scholarly judgment will help test advanced models on difficult real-world problems. You will translate deep subject-matter expertise into precise evaluation material that reveals how well AI systems reason, synthesize evidence, and produce defensible answers.

The Role

OpenTrain AI is seeking an AI Evaluation Question Development Specialist to create rigorous test material for advanced AI model evaluation. You will develop original, high-difficulty question-and-answer pairs at the frontier of your discipline, supported by clear reasoning and authoritative citations.

This is a remote, part-time contractor opportunity with flexible scheduling. Work is completed independently while following project guidelines and quality standards. The expected time commitment is 20 or more hours per week, with compensation ranging from $40 to $90 per hour.

  • Employment type: Contractor and part time
  • Work arrangement: Remote and worldwide
  • Expected commitment: 20+ hours per week
  • Language: English
  • Pay range: $40-$90 per hour

What You'll Do

You will create evaluation content that is challenging, accurate, methodologically sound, and independently verifiable. The work requires careful research, strong analytical judgment, and the ability to improve questions through testing and review.

  • Create original, high-difficulty questions and defensible answers in your area of expertise.
  • Include clear citations, reasoning, and support from primary or authoritative references.
  • Test questions against AI systems and increase complexity when items are not sufficiently challenging.
  • Write with precision and remove ambiguity from questions and answers.
  • Maintain factual accuracy across all evaluation material.
  • Evaluate AI-generated answers for accuracy, challenge level, and ambiguity.
  • Incorporate reviewer feedback while following project quality standards.

Requirements

This opportunity is designed for researchers and domain specialists who can bring scholarly judgment to advanced AI training tasks. You should be able to work independently, assess evidence carefully, and communicate complex ideas in excellent written English.

  • Completed PhD, active PhD candidacy, or equivalent experience as a specialist, researcher, or professor.
  • Demonstrated scholarly research record and strong familiarity with primary literature.
  • Ability to source, triangulate, and cite primary and authoritative references.
  • Advanced analytical thinking and close attention to detail.
  • Skill in composing original questions requiring difficult reasoning and methodological nuance.
  • Ability to evaluate AI-generated answers for accuracy, challenge level, and ambiguity.
  • Excellent written English and reliable independent work habits.

Who Should Apply

This role is a strong fit for PhD researchers, active doctoral candidates, professors, and specialists with equivalent research experience who want to apply their expertise to cutting-edge AI evaluation. You should be comfortable turning deep disciplinary knowledge into original questions that test advanced reasoning rather than basic recall.

Previous AI training or model evaluation experience is helpful but not required. The most important qualifications are scholarly research ability, methodological judgment, precision, and the willingness to strengthen evaluation material through iterative review.

  • Researchers who enjoy developing difficult, intellectually rigorous problems.
  • Subject-matter specialists who can distinguish strong evidence from weak or unsupported claims.
  • Experts interested in shaping how advanced AI systems reason and respond.
  • Independent professionals seeking flexible, remote contractor work in AI training.

How to Apply Through OpenTrain

Create a free OpenTrain account to build your AI training profile and apply for this opportunity. OpenTrain brings together flexible AI training work and helps you develop a credible portfolio as you grow in the field.

Review the role requirements carefully and highlight your research background, scholarly record, subject-matter expertise, and ability to create evidence-based evaluation questions.

  • Create or update your free OpenTrain profile.
  • Showcase relevant research experience and specialist expertise.
  • Apply in minutes through OpenTrain.
  • Complete approved work independently according to project guidelines and quality standards.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

AI Agent Evaluation Scenario Writer

Create realistic evaluation scenarios that test how LLM-based agents handle real-world tasks. This fully remote, part-time contractor role pays $18 to $24 per hour and requires strong English, QA thinking, and basic Python and JavaScript.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Intermediate level
Hourly · $18–$24/hr

Posted Jan 13, 2026

Electrical Engineering AI Evaluation Expert

Use senior electrical engineering judgment to create and evaluate demanding AI tasks across power systems, circuit design, embedded systems, and certification. This remote contract offers 20+ hours per week at $70-$80 per hour.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $70–$80/hr

Posted Jul 29, 2026

Technical Expert Interviewer for AI Evaluation

Conduct structured interviews with senior technical experts in GPU kernels, security research, and ML compilers. This remote U.S. contract role pays $50-$60 per hour for about 10 weekday hours per week.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Entry level
Hourly · $50–$60/hr

Posted Aug 28, 2026