Skip to content
OpenTrain AIFor AI Companies

Web Research Benchmark Designer

Create challenging, objectively verifiable research questions for evaluating frontier AI browsing agents. Use primary sources, precise citations, and structured documentation in a flexible worldwide contract role.

Apply now
OpenTrain AI

Generative AI & RLHF

100% Remote Per task · $30/label

$30/label

Compensation

Worldwide

Eligibility

Intermediate

Experience

Aug 28, 2026

Posted

Open worldwide

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover opportunities, build a professional AI training profile, and apply to projects that let them contribute to the next generation of artificial intelligence.

About AI Training and Benchmark Work

AI training is the human work behind modern artificial intelligence. Contributors create examples, evaluate model behavior, and develop structured tests that help AI systems become more capable, reliable, and useful.

In this role, your research artifacts will support the evaluation of AI browsing agents. You will design difficult investigations that can be checked against authoritative evidence, helping reveal how well models research, reason, and support their answers.

The Role

OpenTrain is hiring a Web Research Benchmark Designer to create challenging research problems for a benchmark focused on frontier AI browsing agents. The work is investigative rather than conventional subject-matter expertise or content writing: you will work backward from verifiable facts to construct questions that remain difficult despite full web access and repeated attempts.

This is a worldwide, part-time contractor opportunity requiring 20 or more hours per week. Compensation is $30 USD per completed label.

  • Role level: Intermediate
  • Work arrangement: Remote and worldwide
  • Engagement: Contractor and part time
  • Language: English
  • Data type: Text
  • Task type: Question answering

What You'll Do

You will produce objective, auditable research artifacts that can be independently validated. Strong performance requires curiosity, persistence, and careful documentation from the first search through the final evidence trail.

  • Create natural-language research questions with short, stable, objectively verifiable answers.
  • Develop independently checkable clues involving dates, people, places, organizations, works, events, records, and quantities.
  • Investigate unfamiliar subjects from scratch using primary sources.
  • Research government and institutional databases, archives, registries, and PDF documents.
  • Record obvious searches performed and the results they returned.
  • Cite exact pages, tables, sections, and other source locations.
  • Deliver structured research outputs with a complete validation trail.
  • Use JSON familiarity when helpful for organized data delivery.

Requirements

You should be able to conduct rigorous open-web research independently and document your reasoning with exceptional sourcing precision. The role requires a master's degree or more than three years of relevant experience, along with native or near-native written English.

Experience with LLM evaluation, red-teaming, or benchmark construction is required. A high tolerance for structured documentation is important because the evidence trail is a central part of every deliverable.

  • Demonstrated open-web research using primary records, institutional databases, archives, registries, and PDF documents.
  • Ability to construct difficult, objectively verifiable research questions from known facts.
  • Precision citing exact pages, tables, sections, and other source locations.
  • Experience with LLM evaluation, red-teaming, or benchmark construction.
  • Strong written English and structured documentation skills.
  • A master's degree or more than three years of relevant experience.

Relevant Backgrounds

This role may suit researchers from a range of investigative and evidence-focused disciplines. The work rewards people who enjoy tracing facts through imperfect information, validating claims, and turning complex findings into clear, reproducible research tasks.

  • Reference librarianship or archival research
  • Special collections research
  • Investigative journalism or professional fact-checking
  • OSINT, due diligence, or KYC
  • Patent or prior-art searching
  • Legal discovery
  • Genealogy
  • Competitive quizzing or puzzle-hunt construction
  • JSON and structured data delivery

Why This Work Matters

AI evaluation depends on carefully designed examples and tests created by people. By building research benchmarks with clear answers and defensible sources, you will help assess whether advanced AI systems can navigate the web, connect evidence, and produce trustworthy results.

OpenTrain offers a way to build experience in a fast-growing AI training industry while applying serious research and source-validation skills to cutting-edge evaluation work.

How to Apply

Create a free OpenTrain account and apply in minutes. Include relevant examples of web research, benchmark construction, LLM evaluation, red-teaming, investigative work, or structured documentation that demonstrate your ability to produce precise, auditable research.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Web Research Benchmark Designer

Create challenging, objectively verifiable research questions for frontier AI browsing agents. This remote U.S. contract offers up to 40 hours weekly for 8 weeks and pays approximately $60 per approved task.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Intermediate level
Per task · $60/label

Posted Sep 1, 2026

Market Research AI Evaluation Expert

Evaluate AI-generated market research and create gold-standard research deliverables as a remote contractor. Earn $30 to $65 per hour while working 20+ hours weekly on future research projects.

Generative AI & RLHF
Text
Remote · Bangladesh, Bhutan, Brazil +14 more
English
Part-time · Flexible
Intermediate level
Hourly · $30–$65/hr

Posted Jul 9, 2026

Architecture And Design Quality Assurance Lead

Review AI-generated architecture and design work, guide remote QA teams, and improve project standards across built-environment topics. This US-based contract role offers up to $75/hour and requires 20+ hours weekly.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Entry level
Hourly · $75/hr

Posted Jul 8, 2026