Create challenging, objectively verifiable research questions for evaluating frontier AI browsing agents. Use primary sources, precise citations, and structured documentation in a flexible worldwide contract role.
About OpenTrain
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover opportunities, build a professional AI training profile, and apply to projects that let them contribute to the next generation of artificial intelligence.
About AI Training and Benchmark Work
AI training is the human work behind modern artificial intelligence. Contributors create examples, evaluate model behavior, and develop structured tests that help AI systems become more capable, reliable, and useful.
In this role, your research artifacts will support the evaluation of AI browsing agents. You will design difficult investigations that can be checked against authoritative evidence, helping reveal how well models research, reason, and support their answers.
The Role
OpenTrain is hiring a Web Research Benchmark Designer to create challenging research problems for a benchmark focused on frontier AI browsing agents. The work is investigative rather than conventional subject-matter expertise or content writing: you will work backward from verifiable facts to construct questions that remain difficult despite full web access and repeated attempts.
This is a worldwide, part-time contractor opportunity requiring 20 or more hours per week. Compensation is $30 USD per completed label.
- Role level: Intermediate
- Work arrangement: Remote and worldwide
- Engagement: Contractor and part time
- Language: English
- Data type: Text
- Task type: Question answering
What You'll Do
You will produce objective, auditable research artifacts that can be independently validated. Strong performance requires curiosity, persistence, and careful documentation from the first search through the final evidence trail.
- Create natural-language research questions with short, stable, objectively verifiable answers.
- Develop independently checkable clues involving dates, people, places, organizations, works, events, records, and quantities.
- Investigate unfamiliar subjects from scratch using primary sources.
- Research government and institutional databases, archives, registries, and PDF documents.
- Record obvious searches performed and the results they returned.
- Cite exact pages, tables, sections, and other source locations.
- Deliver structured research outputs with a complete validation trail.
- Use JSON familiarity when helpful for organized data delivery.
Requirements
You should be able to conduct rigorous open-web research independently and document your reasoning with exceptional sourcing precision. The role requires a master's degree or more than three years of relevant experience, along with native or near-native written English.
Experience with LLM evaluation, red-teaming, or benchmark construction is required. A high tolerance for structured documentation is important because the evidence trail is a central part of every deliverable.
- Demonstrated open-web research using primary records, institutional databases, archives, registries, and PDF documents.
- Ability to construct difficult, objectively verifiable research questions from known facts.
- Precision citing exact pages, tables, sections, and other source locations.
- Experience with LLM evaluation, red-teaming, or benchmark construction.
- Strong written English and structured documentation skills.
- A master's degree or more than three years of relevant experience.
Relevant Backgrounds
This role may suit researchers from a range of investigative and evidence-focused disciplines. The work rewards people who enjoy tracing facts through imperfect information, validating claims, and turning complex findings into clear, reproducible research tasks.
- Reference librarianship or archival research
- Special collections research
- Investigative journalism or professional fact-checking
- OSINT, due diligence, or KYC
- Patent or prior-art searching
- Legal discovery
- Genealogy
- Competitive quizzing or puzzle-hunt construction
- JSON and structured data delivery
Why This Work Matters
AI evaluation depends on carefully designed examples and tests created by people. By building research benchmarks with clear answers and defensible sources, you will help assess whether advanced AI systems can navigate the web, connect evidence, and produce trustworthy results.
OpenTrain offers a way to build experience in a fast-growing AI training industry while applying serious research and source-validation skills to cutting-edge evaluation work.
How to Apply
Create a free OpenTrain account and apply in minutes. Include relevant examples of web research, benchmark construction, LLM evaluation, red-teaming, investigative work, or structured documentation that demonstrate your ability to produce precise, auditable research.