Skip to content
OpenTrain AIFor AI Companies

AI Browsing Benchmark Researcher

Design difficult, objectively verifiable research questions for frontier AI browsing agents. This US-based, part-time contractor role offers 20+ hours per week for experienced web researchers.

OpenTrain AI

Generative AI & RLHF

Remote

1 country

Eligibility

Entry

Experience

Sep 1, 2026

Posted

Open to applicants in

United States

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is hiring an AI Browsing Benchmark Researcher to help advance the systems shaping the future of search and artificial intelligence. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, where contributors can create a profile, discover projects, and apply in minutes.

Creating an OpenTrain account is free, and the platform helps AI-training professionals build a durable portfolio of relevant work and experience.

  • US-based contractor opportunity
  • Part-time engagement with a commitment of 20+ hours per week
  • Apply and manage your AI-training career through OpenTrain

About AI Training and Evaluation

AI training is the human side of building artificial intelligence. People help improve modern systems by creating examples, evaluating model behavior, checking factual accuracy, and testing whether AI can complete demanding tasks reliably.

This role focuses on evaluation research for browsing agents. Your work will help measure whether advanced systems can find, connect, verify, and document information from the open web.

  • Work at the cutting edge of AI development
  • Use human judgment to evaluate advanced model capabilities
  • Contribute to research that helps make AI systems more reliable

The Role

As an AI Browsing Benchmark Researcher, you will design investigative research problems for an evaluation benchmark focused on frontier AI browsing agents. You will work backward from verifiable facts to create questions that remain difficult for advanced systems to answer, even with full web access and multiple attempts.

The work centers on rigorous research, objective verification, and auditable documentation rather than subject-matter instruction or general content writing. Questions should use natural language and have short, stable, objectively verifiable answers.

  • Create challenging research questions for browsing-agent evaluation
  • Develop independently checkable clues across multiple types of information
  • Build evidence trails that another researcher can audit

What You’ll Do

You will investigate unfamiliar subjects from scratch, validate obvious searches, and document the results in a structured format. Your research may connect dates, people, places, organizations, works, events, records, and quantities into questions that require careful browsing and cross-checking.

You will also provide precise citations so each answer can be independently verified. Familiarity with JSON and structured data delivery formats is useful for organizing your work.

  • Write natural-language questions with concise, stable answers
  • Construct difficult questions that remain challenging after multiple search attempts
  • Identify and validate independently checkable clues
  • Research dates, people, places, organizations, works, events, records, and quantities
  • Cite exact pages, tables, sections, and other source locations
  • Deliver findings through disciplined structured documentation

Requirements

You must have demonstrated open-web research ability and confidence investigating unfamiliar topics from scratch. Strong source discipline, native or near-native written English, and a high tolerance for structured documentation are required.

Experience with LLM evaluation, red-teaming, or benchmark construction is required. You must also have a master’s degree or more than three years of professional experience, along with experience in at least one relevant research field or practice.

  • Native or near-native written English
  • Demonstrated open-web research ability
  • Experience with LLM evaluation, red-teaming, or benchmark construction
  • Ability to construct difficult, objectively verifiable research questions
  • Precision when citing exact pages, tables, sections, and other source locations
  • Master’s degree or more than three years of professional experience
  • Experience in reference librarianship, archival research, special collections, investigative journalism, professional fact-checking, OSINT, due diligence, KYC, investigative research, patent or prior-art searching, legal discovery, genealogy, competitive quizzing, or puzzle-hunt construction
  • Familiarity with JSON and structured data delivery formats is useful

Who Should Apply

This opportunity is suited to careful investigators who enjoy tracing evidence across the web and turning complex findings into clear, reproducible research tasks. It may be a strong fit for professionals or specialists from investigative, archival, journalistic, legal, intelligence, research, or puzzle-solving backgrounds.

The listing is marked entry level, while the role requirements call for substantial research or evaluation experience. Applicants should be prepared to demonstrate both rigorous source handling and the ability to document their reasoning in a structured way.

  • Investigative researchers who enjoy unfamiliar subjects and open-ended discovery
  • Fact-checkers, OSINT researchers, journalists, archivists, and reference specialists
  • Professionals experienced in due diligence, KYC, legal discovery, or prior-art research
  • Benchmark builders, red-teamers, and LLM evaluation practitioners
  • Competitive quizzers and puzzle-hunt constructors with strong evidence discipline

How to Apply Through OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. Create a free OpenTrain account, build a profile that reflects your research and AI-evaluation experience, and apply in minutes.

AI training and data-labeling work can provide flexible ways to contribute to cutting-edge technology. OpenTrain helps you present credible experience in one place and discover opportunities that match your skills.

  • Create a free OpenTrain account
  • Highlight relevant research, evaluation, red-teaming, or benchmark experience
  • Show your ability to cite sources and deliver structured documentation
  • Apply for this US-based, part-time contractor opportunity

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

AI Browsing Benchmark Researcher

Create challenging, verifiable research questions for frontier browsing agents using primary records, archives, databases, and precise documentation in a remote 8-week contract.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Intermediate level

Posted Aug 28, 2026

Quantum Optics AI Benchmarking Specialist

Use advanced quantum optics expertise to benchmark AI training work involving interferometry, squeezing, optical loss, and quantum noise. This remote contractor project offers approximately 10 hours per week for 8 to 10 weeks.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $80–$160/hr

Posted Aug 2, 2026

Applied Chemistry Benchmark Specialist

Use advanced chemistry expertise to create and verify rigorous benchmark questions that test whether AI systems genuinely understand complex science. This fully remote freelance role pays $61 to $77 per hour.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $61–$77/hr

Posted Aug 28, 2026