Skip to content
OpenTrain AIFor AI Companies

AI Browsing Benchmark Researcher

Create challenging, verifiable research questions for frontier browsing agents using primary records, archives, databases, and precise documentation in a remote 8-week contract.

OpenTrain AI

Generative AI & RLHF

100% Remote

Worldwide

Eligibility

Intermediate

Experience

Aug 28, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI hires and contracts contributors for projects that help develop modern artificial intelligence, giving professionals a place to build a credible AI-training portfolio and grow their experience.

  • Apply through OpenTrain and manage your AI training work in one place
  • Build a profile that showcases your research and evaluation experience
  • Creating an OpenTrain account is free

About AI Training and Browsing Evaluation

AI training is the human side of building artificial intelligence. People create examples, evaluate model behavior, verify information, and provide structured feedback so AI systems can become more capable and reliable.

Browsing benchmarks help test whether AI agents can investigate questions, follow clues, locate trustworthy evidence, and produce well-supported answers. This role contributes to that process through rigorous web research and benchmark construction.

  • Work on cutting-edge AI evaluation and benchmark development
  • Use human research judgment to test how browsing agents investigate information
  • Contribute to a fast-growing field with remote, flexible contractor opportunities

The Role

OpenTrain is seeking an AI Browsing Benchmark Researcher to create evaluation problems for frontier browsing agents. This contractor assignment combines open-web investigation, research-question construction, source validation, and detailed documentation.

The work focuses on investigative research rather than general content writing or conventional subject-matter consulting. You will investigate unfamiliar subjects from scratch and turn reliable source facts into difficult, objectively verifiable research tasks.

  • Role: AI Browsing Benchmark Researcher
  • Experience level: Intermediate
  • Work type: Contractor and part-time
  • Location: Remote and worldwide
  • Language: Native or near-native written English
  • Contract length: 8 weeks

What You'll Do

You will begin with a stable, objectively verifiable fact and construct a challenging natural-language research question around it. Your work will require independently checkable clues and a clear evidence trail so that another researcher can validate the result.

You will locate primary records across government and institutional databases, archives, registries, and PDF documents. You will then document the searches performed, their results, and the precise evidence supporting each benchmark.

  • Create difficult natural-language research questions from source facts
  • Develop clues involving dates, people, places, organizations, works, events, records, and quantities
  • Investigate unfamiliar topics through open-web research
  • Find and validate primary records in databases, archives, registries, and PDF documents
  • Cite exact pages, tables, and sections
  • Prepare structured validation records documenting searches and results
  • Deliver work in structured formats, with familiarity with JSON considered useful

Requirements

This role requires demonstrated open-web research ability, strong sourcing precision, disciplined structured documentation, and native or near-native written English. Experience with LLM evaluation, red teaming, or benchmark construction is required.

A master's degree or more than three years of experience is required. Relevant investigative and archival experience can include reference librarianship, archival research, special collections, investigative journalism, professional fact-checking, OSINT, due diligence, KYC, patent or prior-art searching, legal discovery, genealogy, competitive quizzing, or puzzle-hunt construction.

  • Demonstrated open-web research using primary records and institutional sources
  • Ability to cite exact pages, tables, and sections with precision
  • Ability to construct objectively verifiable research questions
  • Experience with LLM evaluation, red teaming, or benchmark construction
  • Strong structured documentation skills
  • Master's degree or more than three years of experience
  • Native or near-native written English

Schedule and Compensation

This is a remote contractor assignment for an 8-week contract. The role description specifies a 40-hour workweek, including at least four hours of overlap with Pacific Time. The structured listing indicates a minimum time requirement of 20+ hours per week; applicants should be prepared for the 40-hour schedule described for the assignment.

Compensation details are not specified in the listing.

  • Remote worldwide assignment
  • 8-week contract
  • 40 hours per week according to the role description
  • At least 4 hours of overlap with Pacific Time
  • Structured listing time requirement: 20+ hours per week

Why This Work Matters

Every major AI system depends on people who prepare, review, and evaluate information. By designing demanding browsing benchmarks and validating their evidence, you will help measure whether advanced AI systems can conduct reliable research rather than simply generate plausible-sounding text.

  • Shape how frontier browsing agents are evaluated
  • Apply professional research and verification skills to AI development
  • Build experience in benchmark construction and LLM evaluation
  • Develop a durable portfolio of AI training work through OpenTrain

How to Apply

Create a free OpenTrain account, build your profile, and apply in minutes. Your OpenTrain profile can help you present relevant research, evaluation, and documentation experience as you pursue opportunities in AI training and data labeling.

  • Review the role requirements and schedule
  • Highlight primary-source research and citation experience
  • Showcase LLM evaluation, red teaming, benchmark, archival, or investigative work
  • Apply through OpenTrain

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

AI Browsing Benchmark Researcher

Design difficult, objectively verifiable research questions for frontier AI browsing agents. This US-based, part-time contractor role offers 20+ hours per week for experienced web researchers.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Entry level

Posted Sep 1, 2026

Quantum Optics AI Benchmarking Specialist

Use advanced quantum optics expertise to benchmark AI training work involving interferometry, squeezing, optical loss, and quantum noise. This remote contractor project offers approximately 10 hours per week for 8 to 10 weeks.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $80–$160/hr

Posted Aug 2, 2026

Applied Chemistry Benchmark Specialist

Use advanced chemistry expertise to create and verify rigorous benchmark questions that test whether AI systems genuinely understand complex science. This fully remote freelance role pays $61 to $77 per hour.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $61–$77/hr

Posted Aug 28, 2026