AI Browsing Benchmark Researcher
Design difficult, objectively verifiable research questions for frontier AI browsing agents. This US-based, part-time contractor role offers 20+ hours per week for experienced web researchers.
Posted Sep 1, 2026
Create challenging, verifiable research questions for frontier browsing agents using primary records, archives, databases, and precise documentation in a remote 8-week contract.
Generative AI & RLHF
Worldwide
Eligibility
Intermediate
Experience
Aug 28, 2026
Posted
Open worldwide
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI hires and contracts contributors for projects that help develop modern artificial intelligence, giving professionals a place to build a credible AI-training portfolio and grow their experience.
AI training is the human side of building artificial intelligence. People create examples, evaluate model behavior, verify information, and provide structured feedback so AI systems can become more capable and reliable.
Browsing benchmarks help test whether AI agents can investigate questions, follow clues, locate trustworthy evidence, and produce well-supported answers. This role contributes to that process through rigorous web research and benchmark construction.
OpenTrain is seeking an AI Browsing Benchmark Researcher to create evaluation problems for frontier browsing agents. This contractor assignment combines open-web investigation, research-question construction, source validation, and detailed documentation.
The work focuses on investigative research rather than general content writing or conventional subject-matter consulting. You will investigate unfamiliar subjects from scratch and turn reliable source facts into difficult, objectively verifiable research tasks.
You will begin with a stable, objectively verifiable fact and construct a challenging natural-language research question around it. Your work will require independently checkable clues and a clear evidence trail so that another researcher can validate the result.
You will locate primary records across government and institutional databases, archives, registries, and PDF documents. You will then document the searches performed, their results, and the precise evidence supporting each benchmark.
This role requires demonstrated open-web research ability, strong sourcing precision, disciplined structured documentation, and native or near-native written English. Experience with LLM evaluation, red teaming, or benchmark construction is required.
A master's degree or more than three years of experience is required. Relevant investigative and archival experience can include reference librarianship, archival research, special collections, investigative journalism, professional fact-checking, OSINT, due diligence, KYC, patent or prior-art searching, legal discovery, genealogy, competitive quizzing, or puzzle-hunt construction.
This is a remote contractor assignment for an 8-week contract. The role description specifies a 40-hour workweek, including at least four hours of overlap with Pacific Time. The structured listing indicates a minimum time requirement of 20+ hours per week; applicants should be prepared for the 40-hour schedule described for the assignment.
Compensation details are not specified in the listing.
Every major AI system depends on people who prepare, review, and evaluate information. By designing demanding browsing benchmarks and validating their evidence, you will help measure whether advanced AI systems can conduct reliable research rather than simply generate plausible-sounding text.
Create a free OpenTrain account, build your profile, and apply in minutes. Your OpenTrain profile can help you present relevant research, evaluation, and documentation experience as you pursue opportunities in AI training and data labeling.
Keep exploring
Design difficult, objectively verifiable research questions for frontier AI browsing agents. This US-based, part-time contractor role offers 20+ hours per week for experienced web researchers.
Posted Sep 1, 2026
Use advanced quantum optics expertise to benchmark AI training work involving interferometry, squeezing, optical loss, and quantum noise. This remote contractor project offers approximately 10 hours per week for 8 to 10 weeks.
Posted Aug 2, 2026
Use advanced chemistry expertise to create and verify rigorous benchmark questions that test whether AI systems genuinely understand complex science. This fully remote freelance role pays $61 to $77 per hour.
Posted Aug 28, 2026
Browse related job pages
Expertise
Languages