History and Political Science AI Benchmark Specialist
Create and evaluate rigorous history and political science benchmarks used in AI research. This fully remote contract role offers asynchronous work at $44–$56 per hour for qualified doctoral-level subject-matter experts.
Generative AI & RLHF
100% Remote Hourly · $44–$56/hr
$44–$56/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Aug 15, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the #1 platform for finding and building careers in AI training and data labeling. We help contributors discover specialized projects, build a professional profile, and apply to opportunities in a rapidly growing field where human expertise shapes how advanced AI systems perform.
Free account creation
Remote opportunities across AI training and data labeling
A profile designed to help you build a lasting AI-training portfolio
About AI Benchmark Work
AI training is the human side of building artificial intelligence. Experts create, review, and evaluate examples that help researchers measure model capabilities, identify weaknesses, and improve the quality of AI-generated answers.
In this role, your academic knowledge will contribute to dependable text-based benchmarks that test whether AI systems can reason about complex historical and political subjects rather than rely on surface-level recall.
Contribute to cutting-edge AI research
Use specialized academic expertise in flexible remote work
Help establish reliable standards for evaluating AI capabilities
The Role
OpenTrain AI is recruiting a History and Political Science AI Benchmark Specialist to create and review academic assessment content used in AI research. You will work with rigorous multiple-choice questions covering national security, public policy, business history, environmental history, and Latin American history.
The work requires careful editorial judgment and strong subject-matter expertise. Questions and solutions must be accurate, self-contained, unambiguous, and appropriately challenging for advanced learners.
Contractor and part-time position
Fully remote and asynchronous
Pay range: $44–$56 per hour
English-language work
The description states an expected commitment of 10 or more hours per week; the structured listing indicates 20+ hours per week
What You'll Do
You will help create and maintain gold-standard benchmark materials by authoring new questions, reviewing existing content, and documenting the reasoning behind your decisions. Your work will distinguish genuine conceptual understanding from guessing and surface-level familiarity.
Author original multiple-choice questions that test conceptual understanding
Write one correct answer and nine plausible alternatives for each question
Review questions for accuracy, clarity, completeness, precision, and solvability
Make and explain edits when questions require improvement
Assign medium, hard, or expert difficulty ratings aligned with the intended academic level
Write clear, step-by-step solutions
Support each question with one to five reputable academic references
Evaluate assessment content for rigor and consistency
Help maintain dependable benchmark materials for AI evaluation
Requirements
A PhD or doctoral candidacy is required in History, Political Science, International Relations, or a closely related discipline. Candidates with a master's degree and exceptional depth in a specialized subdomain may also be considered.
You should be able to combine advanced academic knowledge with precise writing and disciplined assessment review. Research publications or policy experience are helpful but are not stated as required.
PhD or doctoral candidacy in History, Political Science, International Relations, or a closely related field
Alternatively, a master's degree with exceptional depth in a specialized subdomain may be considered
Strong command of historiographical methods
Strong command of political theory
Strong command of comparative analysis
Ability to write challenging, unambiguous multiple-choice questions
Ability to assess question solvability, difficulty, accuracy, and rigor
Excellent written English for concise academic explanations and step-by-step solutions
Who Should Apply
This opportunity is suited to historians, political scientists, international relations scholars, and closely related researchers who enjoy translating complex academic ideas into rigorous assessment content. It is listed as entry level, but the required doctoral-level subject expertise makes it especially relevant to doctoral candidates and recently trained specialists.
Academic researchers with expertise in history or political science
Doctoral candidates seeking flexible, remote contract work
International relations specialists with strong comparative-analysis skills
Subject-matter experts who value precision, evidence, and clear reasoning
Writers who can explain complex concepts concisely
How the Work Fits Into AI Training
Modern AI models learn from examples prepared and reviewed by people. By developing questions, answer choices, explanations, and references, you will provide structured human judgment that helps researchers evaluate how well AI systems understand and reason about history and political science.
Create high-quality training and evaluation content
Assess model-relevant questions against academic standards
Apply expert judgment to accuracy, difficulty, and clarity
Use your expertise in politics, governance, elections, public policy, and international relations to evaluate AI responses, design challenging prompts, and improve the accuracy of advanced language models in a remote contractor role.
Lead quality and consistency for political science AI-training projects, reviewing model outputs and trainer work while providing clear, rubric-based feedback. Remote (US), part-time contractor work at up to $70/hr with 20+ hours/week.
Apply advanced historical knowledge to review prompts and AI-generated work for factual accuracy, reasoning quality, and guideline compliance in a remote contract supporting language-model evaluation.