Skip to content
OpenTrain AIFor AI Companies

Evaluation Scenario Writer AI Agent Testing Specialist

Design realistic tests for AI agents, define expected behavior, and review model outputs in a structured writing role. Work worldwide for 20 or more hours per week at $18-$24 per hour.

Apply now
OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $18–$24/hr

$18–$24/hr

Compensation

Worldwide

Eligibility

Intermediate

Experience

Jan 13, 2026

Posted

Open worldwide

The Work

You will create realistic, reusable evaluation scenarios for AI agents powered by large language models. These scenarios simulate real tasks and help measure whether an agent responds and acts as expected.

You will also review agent outputs, improve scenarios for clarity and coverage, and work with developers and other contributors to refine testing frameworks.

  • Design structured evaluation scenarios for LLM-based agents.
  • Define the golden path, meaning the ideal sequence of actions and expected behavior.
  • Describe acceptable alternative behaviors, edge cases, and scoring rules.
  • Annotate task steps and expected outputs in formats such as JSON or YAML.
  • Review agent responses and update scenarios when they are unclear or incomplete.

What It Pays and Takes

This is a part-time contractor role for an intermediate-level contributor. Strong written English and careful, analytical thinking are central to the work.

  • Pay: $18-$24 per hour.
  • Time: 20 or more hours per week.
  • Location: Worldwide.
  • Language: Excellent written English.
  • Required: Basic Python and JavaScript experience.
  • Preferred background: Software testing, quality assurance, data analysis, or NLP annotation.
  • Useful strengths: Scenario design, structured thinking, documentation, and attention to edge cases.
  • Work type: Part-time contract.

How It Works

Apply on OpenTrain. The employer reviews applications there.

About AI Training Work

AI training work uses human-written examples, reviews, and evaluations to improve how artificial intelligence systems behave. People with strong technical or subject knowledge help create clear standards and identify where an AI system succeeds or falls short.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Social Scenario AI Evaluation Generalist

Create realistic social scenarios, assess AI responses, and explain your ratings in clear written feedback. This remote, worldwide project lasts 2-3 weeks, requires 20+ hours weekly, and pays per task.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Intermediate level
Per task

Posted Oct 1, 2026

AI Writing Response Evaluator

Evaluate AI-generated written responses for accuracy, clarity, tone, helpfulness, and instruction following. This remote contract role pays $90-$140 per hour and requires strong English writing or editorial experience.

Generative AI & RLHF
Text
Remote · Andorra, United Arab Emirates, Antigua & Barbuda +227 more
English
Part-time · Flexible
Entry level
Hourly · $90–$140/hr

Posted Sep 19, 2026

Personalized AI Response Evaluation Analyst

Create personal-context prompts, compare AI responses, and explain issues such as unsupported claims or forced connections. This remote, entry-level contract role requires 20+ hours weekly and strong English writing.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Sep 4, 2026