Skip to content
OpenTrain AIFor AI Companies

Web Research Benchmark Designer

Create challenging, objectively verifiable research questions for frontier AI browsing agents. This remote U.S. contract offers up to 40 hours weekly for 8 weeks and pays approximately $60 per approved task.

Apply now
OpenTrain AI

Generative AI & RLHF

Remote Per task · $60/label

$60/label

Compensation

1 country

Eligibility

Intermediate

Experience

Sep 1, 2026

Posted

Open to applicants in

United States

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes.

This project offers the opportunity to turn advanced research skills into practical experience shaping how AI systems browse, investigate, and support their answers.

About AI Training and Browsing Evaluation

AI training is the human work behind modern artificial intelligence. Contributors prepare examples, evaluate model behavior, and provide the detailed feedback that helps AI systems become more accurate and useful.

For browsing-agent evaluation, careful researchers create tests that measure whether an AI can find reliable information, follow evidence trails, and produce answers supported by objectively verifiable sources.

The Web Research Benchmark Designer Role

As a Web Research Benchmark Designer, you will create difficult, auditable research problems for a benchmark focused on frontier AI browsing agents. You will work backward from verifiable facts to design questions that remain challenging even when an AI system has full web access and can make repeated attempts.

The role combines investigative research, source validation, structured documentation, and evaluation of whether an AI system can locate and support an objectively verifiable answer.

  • Location: United States
  • Project length: Approximately 8 weeks
  • Availability: 20 or more hours per week, with up to 40 hours available
  • Engagement: Remote contractor and part-time project
  • Compensation: Approximately $60 per approved task, paid in USD

What You'll Do

You will research unfamiliar subjects from scratch and turn reliable findings into structured benchmark tasks. Your work should give evaluators a clear way to determine whether an AI browsing agent reached a correct answer and supported it with strong evidence.

  • Design natural-language research questions with short, stable, objectively verifiable answers.
  • Create independently checkable clues involving dates, people, places, organizations, works, events, records, and quantities.
  • Locate primary records through government and institutional databases, archives, registries, and PDF documents.
  • Build complete evidence trails with exact pages, tables, and sections.
  • Record searches and validation results in a structured format.
  • Evaluate whether research questions and supporting evidence are precise, auditable, and appropriately difficult.

Required Experience and Skills

This is an intermediate-level research role for contributors who can independently investigate complex questions and distinguish primary evidence from general web pages. Strong written English and careful documentation are essential.

  • Demonstrated open-web research using primary records, archives, registries, institutional databases, and PDF documents.
  • Experience in investigative research, archival research, OSINT, fact-checking, due diligence, legal discovery, genealogy, records research, or a comparable field.
  • Ability to construct difficult, objectively verifiable research questions and document precise supporting evidence.
  • Experience with LLM evaluation, red-teaming, benchmark construction, or related AI quality work is valuable.
  • Native or near-native written English.
  • Master's degree or more than three years of professional experience.
  • Comfort with detailed documentation.
  • Familiarity with JSON and structured data delivery formats is helpful.

Who Should Apply

This project may suit researchers who enjoy tracing facts to original records, testing claims, and building transparent evidence trails. It is especially relevant for people with a background in investigative, archival, OSINT, fact-checking, legal-discovery, genealogy, or comparable records research.

The work is fully remote and designed for contributors located in the United States who can commit at least 20 hours per week during the project.

  • Experienced open-web investigators who prioritize source quality and precision.
  • Researchers comfortable navigating institutional databases, archives, registries, and scanned documents.
  • AI evaluators and red-teamers interested in testing browsing-agent capabilities.
  • Detail-oriented contributors who can deliver consistent, structured documentation.

How to Apply Through OpenTrain

Create or use your free OpenTrain account to apply for this project. OpenTrain helps contributors build a credible AI training portfolio while finding work that matches their research and evaluation experience.

If selected, you will complete the project remotely as a contractor, following the assignment requirements and submitting tasks for approval and payment.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Web Research Benchmark Designer

Create challenging, objectively verifiable research questions for evaluating frontier AI browsing agents. Use primary sources, precise citations, and structured documentation in a flexible worldwide contract role.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Intermediate level
Per task · $30/label

Posted Aug 28, 2026

Market Research AI Evaluation Expert

Evaluate AI-generated market research and create gold-standard research deliverables as a remote contractor. Earn $30 to $65 per hour while working 20+ hours weekly on future research projects.

Generative AI & RLHF
Text
Remote · Bangladesh, Bhutan, Brazil +14 more
English
Part-time · Flexible
Intermediate level
Hourly · $30–$65/hr

Posted Jul 9, 2026

Architecture And Design Quality Assurance Lead

Review AI-generated architecture and design work, guide remote QA teams, and improve project standards across built-environment topics. This US-based contract role offers up to $75/hour and requires 20+ hours weekly.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Entry level
Hourly · $75/hr

Posted Jul 8, 2026