Skip to content
OpenTrain AIFor AI Companies

History and Political Science AI Benchmark Specialist

Create and evaluate rigorous history and political science benchmarks used in AI research. This fully remote contract role offers asynchronous work at $44–$56 per hour for qualified doctoral-level subject-matter experts.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $44–$56/hr

$44–$56/hr

Compensation

Worldwide

Eligibility

Entry

Experience

Aug 15, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the #1 platform for finding and building careers in AI training and data labeling. We help contributors discover specialized projects, build a professional profile, and apply to opportunities in a rapidly growing field where human expertise shapes how advanced AI systems perform.

  • Free account creation
  • Remote opportunities across AI training and data labeling
  • A profile designed to help you build a lasting AI-training portfolio

About AI Benchmark Work

AI training is the human side of building artificial intelligence. Experts create, review, and evaluate examples that help researchers measure model capabilities, identify weaknesses, and improve the quality of AI-generated answers.

In this role, your academic knowledge will contribute to dependable text-based benchmarks that test whether AI systems can reason about complex historical and political subjects rather than rely on surface-level recall.

  • Contribute to cutting-edge AI research
  • Use specialized academic expertise in flexible remote work
  • Help establish reliable standards for evaluating AI capabilities

The Role

OpenTrain AI is recruiting a History and Political Science AI Benchmark Specialist to create and review academic assessment content used in AI research. You will work with rigorous multiple-choice questions covering national security, public policy, business history, environmental history, and Latin American history.

The work requires careful editorial judgment and strong subject-matter expertise. Questions and solutions must be accurate, self-contained, unambiguous, and appropriately challenging for advanced learners.

  • Contractor and part-time position
  • Fully remote and asynchronous
  • Pay range: $44–$56 per hour
  • English-language work
  • The description states an expected commitment of 10 or more hours per week; the structured listing indicates 20+ hours per week

What You'll Do

You will help create and maintain gold-standard benchmark materials by authoring new questions, reviewing existing content, and documenting the reasoning behind your decisions. Your work will distinguish genuine conceptual understanding from guessing and surface-level familiarity.

  • Author original multiple-choice questions that test conceptual understanding
  • Write one correct answer and nine plausible alternatives for each question
  • Review questions for accuracy, clarity, completeness, precision, and solvability
  • Make and explain edits when questions require improvement
  • Assign medium, hard, or expert difficulty ratings aligned with the intended academic level
  • Write clear, step-by-step solutions
  • Support each question with one to five reputable academic references
  • Evaluate assessment content for rigor and consistency
  • Help maintain dependable benchmark materials for AI evaluation

Requirements

A PhD or doctoral candidacy is required in History, Political Science, International Relations, or a closely related discipline. Candidates with a master's degree and exceptional depth in a specialized subdomain may also be considered.

You should be able to combine advanced academic knowledge with precise writing and disciplined assessment review. Research publications or policy experience are helpful but are not stated as required.

  • PhD or doctoral candidacy in History, Political Science, International Relations, or a closely related field
  • Alternatively, a master's degree with exceptional depth in a specialized subdomain may be considered
  • Strong command of historiographical methods
  • Strong command of political theory
  • Strong command of comparative analysis
  • Ability to write challenging, unambiguous multiple-choice questions
  • Ability to assess question solvability, difficulty, accuracy, and rigor
  • Excellent written English for concise academic explanations and step-by-step solutions

Who Should Apply

This opportunity is suited to historians, political scientists, international relations scholars, and closely related researchers who enjoy translating complex academic ideas into rigorous assessment content. It is listed as entry level, but the required doctoral-level subject expertise makes it especially relevant to doctoral candidates and recently trained specialists.

  • Academic researchers with expertise in history or political science
  • Doctoral candidates seeking flexible, remote contract work
  • International relations specialists with strong comparative-analysis skills
  • Subject-matter experts who value precision, evidence, and clear reasoning
  • Writers who can explain complex concepts concisely

How the Work Fits Into AI Training

Modern AI models learn from examples prepared and reviewed by people. By developing questions, answer choices, explanations, and references, you will provide structured human judgment that helps researchers evaluate how well AI systems understand and reason about history and political science.

  • Create high-quality training and evaluation content
  • Assess model-relevant questions against academic standards
  • Apply expert judgment to accuracy, difficulty, and clarity
  • Work remotely with an asynchronous schedule

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all jobs

Political Science AI Evaluation Expert

Use your expertise in politics, governance, elections, public policy, and international relations to evaluate AI responses, design challenging prompts, and improve the accuracy of advanced language models in a remote contractor role.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 15, 2026

Political Science Quality Assurance Lead

Lead quality and consistency for political science AI-training projects, reviewing model outputs and trainer work while providing clear, rubric-based feedback. Remote (US), part-time contractor work at up to $70/hr with 20+ hours/week.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Entry level
Hourly · $70/hr

Posted Jul 9, 2026

History Domain Reviewer for AI Model Evaluation

Apply advanced historical knowledge to review prompts and AI-generated work for factual accuracy, reasoning quality, and guideline compliance in a remote contract supporting language-model evaluation.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 15, 2026