Create challenging, objectively verifiable research questions for frontier AI browsing agents. This remote U.S. contract offers up to 40 hours weekly for 8 weeks and pays approximately $60 per approved task.
About OpenTrain
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes.
This project offers the opportunity to turn advanced research skills into practical experience shaping how AI systems browse, investigate, and support their answers.
About AI Training and Browsing Evaluation
AI training is the human work behind modern artificial intelligence. Contributors prepare examples, evaluate model behavior, and provide the detailed feedback that helps AI systems become more accurate and useful.
For browsing-agent evaluation, careful researchers create tests that measure whether an AI can find reliable information, follow evidence trails, and produce answers supported by objectively verifiable sources.
The Web Research Benchmark Designer Role
As a Web Research Benchmark Designer, you will create difficult, auditable research problems for a benchmark focused on frontier AI browsing agents. You will work backward from verifiable facts to design questions that remain challenging even when an AI system has full web access and can make repeated attempts.
The role combines investigative research, source validation, structured documentation, and evaluation of whether an AI system can locate and support an objectively verifiable answer.
- Location: United States
- Project length: Approximately 8 weeks
- Availability: 20 or more hours per week, with up to 40 hours available
- Engagement: Remote contractor and part-time project
- Compensation: Approximately $60 per approved task, paid in USD
What You'll Do
You will research unfamiliar subjects from scratch and turn reliable findings into structured benchmark tasks. Your work should give evaluators a clear way to determine whether an AI browsing agent reached a correct answer and supported it with strong evidence.
- Design natural-language research questions with short, stable, objectively verifiable answers.
- Create independently checkable clues involving dates, people, places, organizations, works, events, records, and quantities.
- Locate primary records through government and institutional databases, archives, registries, and PDF documents.
- Build complete evidence trails with exact pages, tables, and sections.
- Record searches and validation results in a structured format.
- Evaluate whether research questions and supporting evidence are precise, auditable, and appropriately difficult.
Required Experience and Skills
This is an intermediate-level research role for contributors who can independently investigate complex questions and distinguish primary evidence from general web pages. Strong written English and careful documentation are essential.
- Demonstrated open-web research using primary records, archives, registries, institutional databases, and PDF documents.
- Experience in investigative research, archival research, OSINT, fact-checking, due diligence, legal discovery, genealogy, records research, or a comparable field.
- Ability to construct difficult, objectively verifiable research questions and document precise supporting evidence.
- Experience with LLM evaluation, red-teaming, benchmark construction, or related AI quality work is valuable.
- Native or near-native written English.
- Master's degree or more than three years of professional experience.
- Comfort with detailed documentation.
- Familiarity with JSON and structured data delivery formats is helpful.
Who Should Apply
This project may suit researchers who enjoy tracing facts to original records, testing claims, and building transparent evidence trails. It is especially relevant for people with a background in investigative, archival, OSINT, fact-checking, legal-discovery, genealogy, or comparable records research.
The work is fully remote and designed for contributors located in the United States who can commit at least 20 hours per week during the project.
- Experienced open-web investigators who prioritize source quality and precision.
- Researchers comfortable navigating institutional databases, archives, registries, and scanned documents.
- AI evaluators and red-teamers interested in testing browsing-agent capabilities.
- Detail-oriented contributors who can deliver consistent, structured documentation.
How to Apply Through OpenTrain
Create or use your free OpenTrain account to apply for this project. OpenTrain helps contributors build a credible AI training portfolio while finding work that matches their research and evaluation experience.
If selected, you will complete the project remotely as a contractor, following the assignment requirements and submitting tasks for approval and payment.