Skip to content
OpenTrain AIFor AI Companies

English LLM Evaluation Generalist

Evaluate ChatGPT and Claude responses, create challenging prompts, and explain model strengths and weaknesses in clear American English. Work remotely for 20+ hours per week at $20-$30 per hour.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $20–$30/hr

$20–$30/hr

Compensation

Worldwide

Eligibility

Entry

Experience

Jul 15, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring and contracting contributors for practical projects that help improve how modern artificial intelligence systems understand and generate language.

  • Build experience in a fast-growing AI training industry
  • Work remotely from anywhere in the world
  • Create a free OpenTrain account and apply in minutes

About AI Model Evaluation

AI models learn from human-created examples, comparisons, ratings, and feedback. In evaluation work, contributors assess model responses against clear guidelines and provide careful judgments that help make AI systems more accurate, useful, and consistent.

  • Support the development of generative AI through human feedback
  • Use critical thinking and communication skills in hands-on model review
  • Work part time with a schedule of 20 or more hours per week

The Role

OpenTrain is seeking an English LLM Evaluation Generalist to create prompts, compare large language model responses, and provide structured verbal feedback. You will work across a wide range of topics, evaluate outputs from ChatGPT and Claude in real time, and record your screen and microphone during each session.

This entry-level contractor role is suited to someone with strong American English communication, practical experience using generative AI tools, and the judgment to explain model strengths, weaknesses, and meaningful differences.

  • Employment type: Part-time contractor
  • Experience level: Entry level
  • Time requirement: 20+ hours per week
  • Pay: $20-$30 per hour
  • Location: Worldwide and fully remote
  • Language: Fluent American English

What You'll Do

You will evaluate language model behavior through prompt creation, side-by-side comparison, and detailed feedback. Each session requires focused work in a distraction-free setting, consistent use of evaluation guidance, and clear communication of your reasoning.

  • Create prompts that challenge large language models across varied topics
  • Compare responses from ChatGPT and Claude
  • Identify strengths, weaknesses, and meaningful differences between responses
  • Record your screen and microphone while completing evaluation sessions
  • Deliver detailed verbal feedback
  • Interpret guidance documentation and apply consistent evaluation standards
  • Maintain clarity, professionalism, and accuracy throughout each session

Requirements

Applicants should be fluent in spoken and written American English and comfortable explaining preferences and judgments clearly. You should also be able to work reliably with the required recording setup and apply detailed instructions consistently.

  • High-level fluency in spoken and written American English
  • Experience using ChatGPT, Claude, Gemini, or similar large language models
  • Familiarity with AI evaluation, content evaluation, data annotation, or quality assurance
  • Strong critical thinking and analytical reasoning
  • Excellent attention to detail
  • Reliable computer and stable internet connection
  • Ability to record high-quality audio and video without issues
  • Comfort comparing ChatGPT and Claude responses
  • Ability to explain model strengths, weaknesses, and preferences clearly

Helpful Background

Research experience or other work involving structured comparison and written or verbal analysis can help you succeed in this role. The work centers on practical judgment, careful reasoning, and consistent communication.

  • Research experience
  • Structured comparison or analysis experience
  • Clear written and verbal communication

How to Apply Through OpenTrain

Create a free OpenTrain account to build your AI training profile and apply for this opportunity. OpenTrain brings contributors into a growing field where human reviewers help shape the behavior of cutting-edge AI systems.

  • Review the role requirements
  • Create or update your OpenTrain profile
  • Apply through OpenTrain in minutes
  • Prepare a reliable computer, internet connection, and recording setup

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

LLM Evaluation and AI Data Analyst

Help improve large language models by evaluating AI-generated responses, researching claims, analyzing data, and writing clear feedback. This entry-level contractor role offers remote, part-time work of 20+ hours per week.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 25, 2026

Biology LLM Evaluation Expert

Help improve large language models by creating challenging biology problems, writing rigorous solutions, and evaluating model reasoning from undergraduate through PhD level. This remote expert contract requires 20+ hours weekly.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Expert level

Posted Jul 17, 2026

Sports LLM Evaluation Expert

Use deep sports knowledge to write challenging prompts, evaluate large language model responses, and identify factual or reasoning issues. This remote contractor role offers 20+ hours per week for experts with a master's degree and three years of relevant experience.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 11, 2026