AI Browsing Benchmark Researcher
Create challenging, verifiable research questions for frontier browsing agents using primary records, archives, databases, and precise documentation in a remote 8-week contract.
Posted Aug 28, 2026
Design difficult, objectively verifiable research questions for frontier AI browsing agents. This US-based, part-time contractor role offers 20+ hours per week for experienced web researchers.
Generative AI & RLHF
1 country
Eligibility
Entry
Experience
Sep 1, 2026
Posted
Open to applicants in
OpenTrain AI is hiring an AI Browsing Benchmark Researcher to help advance the systems shaping the future of search and artificial intelligence. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, where contributors can create a profile, discover projects, and apply in minutes.
Creating an OpenTrain account is free, and the platform helps AI-training professionals build a durable portfolio of relevant work and experience.
AI training is the human side of building artificial intelligence. People help improve modern systems by creating examples, evaluating model behavior, checking factual accuracy, and testing whether AI can complete demanding tasks reliably.
This role focuses on evaluation research for browsing agents. Your work will help measure whether advanced systems can find, connect, verify, and document information from the open web.
As an AI Browsing Benchmark Researcher, you will design investigative research problems for an evaluation benchmark focused on frontier AI browsing agents. You will work backward from verifiable facts to create questions that remain difficult for advanced systems to answer, even with full web access and multiple attempts.
The work centers on rigorous research, objective verification, and auditable documentation rather than subject-matter instruction or general content writing. Questions should use natural language and have short, stable, objectively verifiable answers.
You will investigate unfamiliar subjects from scratch, validate obvious searches, and document the results in a structured format. Your research may connect dates, people, places, organizations, works, events, records, and quantities into questions that require careful browsing and cross-checking.
You will also provide precise citations so each answer can be independently verified. Familiarity with JSON and structured data delivery formats is useful for organizing your work.
You must have demonstrated open-web research ability and confidence investigating unfamiliar topics from scratch. Strong source discipline, native or near-native written English, and a high tolerance for structured documentation are required.
Experience with LLM evaluation, red-teaming, or benchmark construction is required. You must also have a master’s degree or more than three years of professional experience, along with experience in at least one relevant research field or practice.
This opportunity is suited to careful investigators who enjoy tracing evidence across the web and turning complex findings into clear, reproducible research tasks. It may be a strong fit for professionals or specialists from investigative, archival, journalistic, legal, intelligence, research, or puzzle-solving backgrounds.
The listing is marked entry level, while the role requirements call for substantial research or evaluation experience. Applicants should be prepared to demonstrate both rigorous source handling and the ability to document their reasoning in a structured way.
OpenTrain AI is the hiring and contracting organization for this role. Create a free OpenTrain account, build a profile that reflects your research and AI-evaluation experience, and apply in minutes.
AI training and data-labeling work can provide flexible ways to contribute to cutting-edge technology. OpenTrain helps you present credible experience in one place and discover opportunities that match your skills.
Keep exploring
Create challenging, verifiable research questions for frontier browsing agents using primary records, archives, databases, and precise documentation in a remote 8-week contract.
Posted Aug 28, 2026
Use advanced quantum optics expertise to benchmark AI training work involving interferometry, squeezing, optical loss, and quantum noise. This remote contractor project offers approximately 10 hours per week for 8 to 10 weeks.
Posted Aug 2, 2026
Use advanced chemistry expertise to create and verify rigorous benchmark questions that test whether AI systems genuinely understand complex science. This fully remote freelance role pays $61 to $77 per hour.
Posted Aug 28, 2026
Browse related job pages
Expertise
Locations
Languages