Create realistic evaluation scenarios that test how LLM-based agents handle real-world tasks. This fully remote, part-time contractor role pays $18 to $24 per hour and requires strong English, QA thinking, and basic Python and JavaScript.
Generative AI & RLHF
100% Remote Hourly · $18–$24/hr
$18–$24/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Jan 13, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover opportunities to teach and evaluate artificial intelligence.
This contractor position is fully remote and open worldwide. The role requires a reliable laptop, stable internet connection, and consistent availability.
Contractor and part-time engagement
Worldwide remote opportunity
20+ hours per week
Compensation of $18 to $24 per hour
About AI Agent Evaluation Work
AI training is the human side of building modern artificial intelligence. People create examples, review model behavior, and provide structured feedback that helps AI systems become more accurate, useful, and reliable.
In this role, you will help evaluate LLM-based agents by designing realistic tasks, defining expected behavior, and documenting how outputs should be judged. Your work will support the testing and refinement of systems that perform complex, real-world activities.
Work directly with AI-generated outputs and agent logs
Turn real-world tasks into structured, reusable evaluations
Help improve the clarity and coverage of AI testing frameworks
The Role
OpenTrain AI is seeking an analytical Evaluation Scenario Writer and AI Agent Testing Specialist with strong QA-style thinking and excellent written English. You will create reproducible evaluation scenarios for LLM-based agents and establish the standards used to assess their performance.
This intermediate-level role is a strong fit for someone with experience in software testing, QA, test case design, data analysis, or NLP annotation. You should be comfortable working with structured formats such as JSON or YAML and reading or editing simple Python and JavaScript scripts.
Subject matter: LLM agent testing and evaluation design
Experience level: Intermediate
Language: English
Data type: Text
Labeling focus: Evaluation rating
What You'll Do
You will design realistic, reusable scenarios that simulate real-world tasks for LLM-based agents. Each scenario should clearly communicate the task, expected outcomes, acceptable variations, and conditions for evaluating agent performance.
You will also review agent outputs, identify gaps in scenario coverage, and refine evaluation materials for clarity and consistency. The work involves collaborating with developers and other contributors to test and improve evaluation frameworks.
Design structured evaluation scenarios for LLM-based agents
Define the golden path and gold-standard agent behavior
Annotate task steps and expected outputs
Document edge cases and scoring logic
Specify acceptable variations in agent behavior
Review agent outputs and agent logs
Iterate on scenarios to improve clarity and coverage
Collaborate with developers and other contributors
Requirements
Applicants should be able to follow complex guidelines accurately, switch between topics quickly, and produce documentation that is clear, precise, and unambiguous. A background in QA, software testing, data analysis, or NLP annotation is strongly preferred.
A bachelor's and/or master's degree in computer science, software engineering, data science or analytics, artificial intelligence or machine learning, computational linguistics or NLP, information systems, or a related field is required. Basic working experience with Python and JavaScript is also required.
Bachelor's and/or master's degree in a relevant technical or analytical field
Prior experience in QA, software testing, test case design, data analysis, or NLP annotation
Ability to design reproducible test scenarios with strong coverage and edge cases
Comfort using JSON and/or YAML to describe scenarios
Ability to define gold-standard behavior, acceptable variations, and scoring logic
Basic Python and JavaScript experience for reading or editing simple scripts
Strong written English skills
Comfort working with AI-generated outputs, agent logs, and prompt-based behaviors
Reliable laptop, stable internet connection, and consistent availability
Why Join AI Training Work
AI training and data labeling are growing fields that offer remote ways to contribute to cutting-edge technology. Contributors apply analytical, language, and technical skills to the examples and evaluations used to shape how AI systems behave.
Through OpenTrain, you can build experience in an industry where human judgment remains essential to testing and improving modern AI. This opportunity offers part-time flexibility while engaging with advanced LLM evaluation work.
Fully remote work from anywhere worldwide
Part-time schedule of 20+ hours per week
Hands-on experience with LLM agent evaluation
Opportunity to apply QA, testing, writing, and technical skills
Use advanced research expertise to create challenging, evidence-based questions that evaluate how AI models reason and synthesize information. This remote contractor role pays $40-$90 per hour and requires 20+ hours weekly.
Use hands-on social science research expertise to design rigorous evaluation tasks for frontier AI benchmarking. Create surveys, coded datasets, statistical analyses, research artifacts, and detailed grading rubrics as a remote contractor.
Use your fantasy sports expertise to evaluate AI outputs, validate scoring and playoff logic, and calibrate draft boards and projection models. This worldwide, part-time contractor role requires 20+ hours weekly.