Skip to content
OpenTrain AIFor AI Companies

AI Agent Evaluation Scenario Writer

Create realistic evaluation scenarios that test how LLM-based agents handle real-world tasks. This fully remote, part-time contractor role pays $18 to $24 per hour and requires strong English, QA thinking, and basic Python and JavaScript.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $18–$24/hr

$18–$24/hr

Compensation

Worldwide

Eligibility

Intermediate

Experience

Jan 13, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover opportunities to teach and evaluate artificial intelligence.

This contractor position is fully remote and open worldwide. The role requires a reliable laptop, stable internet connection, and consistent availability.

  • Contractor and part-time engagement
  • Worldwide remote opportunity
  • 20+ hours per week
  • Compensation of $18 to $24 per hour

About AI Agent Evaluation Work

AI training is the human side of building modern artificial intelligence. People create examples, review model behavior, and provide structured feedback that helps AI systems become more accurate, useful, and reliable.

In this role, you will help evaluate LLM-based agents by designing realistic tasks, defining expected behavior, and documenting how outputs should be judged. Your work will support the testing and refinement of systems that perform complex, real-world activities.

  • Work directly with AI-generated outputs and agent logs
  • Turn real-world tasks into structured, reusable evaluations
  • Help improve the clarity and coverage of AI testing frameworks

The Role

OpenTrain AI is seeking an analytical Evaluation Scenario Writer and AI Agent Testing Specialist with strong QA-style thinking and excellent written English. You will create reproducible evaluation scenarios for LLM-based agents and establish the standards used to assess their performance.

This intermediate-level role is a strong fit for someone with experience in software testing, QA, test case design, data analysis, or NLP annotation. You should be comfortable working with structured formats such as JSON or YAML and reading or editing simple Python and JavaScript scripts.

  • Subject matter: LLM agent testing and evaluation design
  • Experience level: Intermediate
  • Language: English
  • Data type: Text
  • Labeling focus: Evaluation rating

What You'll Do

You will design realistic, reusable scenarios that simulate real-world tasks for LLM-based agents. Each scenario should clearly communicate the task, expected outcomes, acceptable variations, and conditions for evaluating agent performance.

You will also review agent outputs, identify gaps in scenario coverage, and refine evaluation materials for clarity and consistency. The work involves collaborating with developers and other contributors to test and improve evaluation frameworks.

  • Design structured evaluation scenarios for LLM-based agents
  • Define the golden path and gold-standard agent behavior
  • Annotate task steps and expected outputs
  • Document edge cases and scoring logic
  • Specify acceptable variations in agent behavior
  • Review agent outputs and agent logs
  • Iterate on scenarios to improve clarity and coverage
  • Collaborate with developers and other contributors

Requirements

Applicants should be able to follow complex guidelines accurately, switch between topics quickly, and produce documentation that is clear, precise, and unambiguous. A background in QA, software testing, data analysis, or NLP annotation is strongly preferred.

A bachelor's and/or master's degree in computer science, software engineering, data science or analytics, artificial intelligence or machine learning, computational linguistics or NLP, information systems, or a related field is required. Basic working experience with Python and JavaScript is also required.

  • Bachelor's and/or master's degree in a relevant technical or analytical field
  • Prior experience in QA, software testing, test case design, data analysis, or NLP annotation
  • Ability to design reproducible test scenarios with strong coverage and edge cases
  • Comfort using JSON and/or YAML to describe scenarios
  • Ability to define gold-standard behavior, acceptable variations, and scoring logic
  • Basic Python and JavaScript experience for reading or editing simple scripts
  • Strong written English skills
  • Comfort working with AI-generated outputs, agent logs, and prompt-based behaviors
  • Reliable laptop, stable internet connection, and consistent availability

Why Join AI Training Work

AI training and data labeling are growing fields that offer remote ways to contribute to cutting-edge technology. Contributors apply analytical, language, and technical skills to the examples and evaluations used to shape how AI systems behave.

Through OpenTrain, you can build experience in an industry where human judgment remains essential to testing and improving modern AI. This opportunity offers part-time flexibility while engaging with advanced LLM evaluation work.

  • Fully remote work from anywhere worldwide
  • Part-time schedule of 20+ hours per week
  • Hands-on experience with LLM agent evaluation
  • Opportunity to apply QA, testing, writing, and technical skills

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

AI Evaluation Question Development Specialist

Use advanced research expertise to create challenging, evidence-based questions that evaluate how AI models reason and synthesize information. This remote contractor role pays $40-$90 per hour and requires 20+ hours weekly.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $40–$90/hr

Posted Sep 8, 2026

Social Science AI Evaluation Researcher

Use hands-on social science research expertise to design rigorous evaluation tasks for frontier AI benchmarking. Create surveys, coded datasets, statistical analyses, research artifacts, and detailed grading rubrics as a remote contractor.

Generative AI & RLHF
Document
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $30–$50/hr

Posted Aug 13, 2026

Fantasy Sports AI Evaluation Expert

Use your fantasy sports expertise to evaluate AI outputs, validate scoring and playoff logic, and calibrate draft boards and projection models. This worldwide, part-time contractor role requires 20+ hours weekly.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Sep 6, 2026