Design realistic tests for AI agents, define expected behavior, and review model outputs in a structured writing role. Work worldwide for 20 or more hours per week at $18-$24 per hour.
The Work
You will create realistic, reusable evaluation scenarios for AI agents powered by large language models. These scenarios simulate real tasks and help measure whether an agent responds and acts as expected.
You will also review agent outputs, improve scenarios for clarity and coverage, and work with developers and other contributors to refine testing frameworks.
- Design structured evaluation scenarios for LLM-based agents.
- Define the golden path, meaning the ideal sequence of actions and expected behavior.
- Describe acceptable alternative behaviors, edge cases, and scoring rules.
- Annotate task steps and expected outputs in formats such as JSON or YAML.
- Review agent responses and update scenarios when they are unclear or incomplete.
What It Pays and Takes
This is a part-time contractor role for an intermediate-level contributor. Strong written English and careful, analytical thinking are central to the work.
- Pay: $18-$24 per hour.
- Time: 20 or more hours per week.
- Location: Worldwide.
- Language: Excellent written English.
- Required: Basic Python and JavaScript experience.
- Preferred background: Software testing, quality assurance, data analysis, or NLP annotation.
- Useful strengths: Scenario design, structured thinking, documentation, and attention to edge cases.
- Work type: Part-time contract.
How It Works
Apply on OpenTrain. The employer reviews applications there.
About AI Training Work
AI training work uses human-written examples, reviews, and evaluations to improve how artificial intelligence systems behave. People with strong technical or subject knowledge help create clear standards and identify where an AI system succeeds or falls short.