Skip to content
OpenTrain AIFor AI Companies

English LLM Evaluation Generalist

Join OpenTrain to evaluate large language model outputs, create challenging prompts, and deliver recorded verbal feedback; remote, contract role (20+ hrs/week) paying $20–$30/hr. Entry-level friendly for strong American English speakers with LLM experience.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $20–$30/hr

$20–$30/hr

Compensation

Worldwide

Eligibility

Entry

Experience

Jul 15, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help freelancers discover specialized AI training work, build a unified portfolio, and grow durable freelance careers in an expanding field.

This role is hired and contracted through OpenTrain AI — you will work directly with our team on model-evaluation projects and have your contributions tracked in your OpenTrain profile.

About AI training work

AI training (also called data labeling or human feedback work) is the human side of building modern AI: people create, evaluate, and refine examples that teach models how to behave. These projects are mostly remote, flexible, and accessible to contributors with clear communication and attention to detail.

As an evaluator you will directly shape how state-of-the-art language models perform by crafting prompts, comparing outputs, and providing structured feedback.

The role

You will serve as an English LLM Evaluation Generalist focused on prompt creation, side-by-side model comparison, and structured verbal feedback. Work is text-focused and will include evaluation rating and text-generation tasks.

This is a contract, part-time role expecting 20+ hours per week; pay is $20–$30 per hour. The position is entry level but requires fluency in American English and familiarity with large language models.

  • Work type: Contractor, Part-time
  • Time commitment: 20+ hours/week
  • Pay: USD $20–$30 per hour
  • Data type: Text; label types: EVALUATION_RATING, TEXT_GENERATION

What you'll do

You will create diverse prompts, run them through LLMs, compare outputs side-by-side, and record detailed spoken evaluations while capturing your screen and microphone.

Accuracy, clear reasoning, and consistent application of evaluation guidelines are essential — you will explain strengths, weaknesses, and meaningful differences between model responses.

  • Write prompts that test models across varied topics and styles.
  • Compare responses from major LLMs (for example, ChatGPT and Claude) and highlight differences in quality and behavior.
  • Record high-quality screen and audio during sessions and deliver structured verbal feedback in a distraction-free environment.
  • Apply provided evaluation guidelines consistently and document your judgments clearly.

Requirements

You must meet all core requirements below to perform the role reliably and professionally.

Technical readiness — a reliable computer, stable internet, and the ability to record high-quality audio and video — is non-negotiable because sessions are recorded.

  • Fluent spoken and written American English.
  • Experience using large language models such as ChatGPT, Claude, Gemini, or similar tools.
  • Familiarity with AI evaluation tasks, data annotation, QA, or content evaluation.
  • Strong critical thinking, analytical reasoning, and attention to detail.
  • Reliable computer, stable internet connection, and the ability to record screen and microphone with clear audio.
  • Comfort explaining model strengths, weaknesses, and preferences clearly in spoken feedback.

Helpful background

You don't need deep research credentials, but experience that demonstrates structured comparison and clear analysis will help you succeed.

Relevant experience examples include prior AI evaluation, QA work, research tasks, or roles requiring systematic written or verbal analysis.

  • Research experience or structured analysis roles are a plus.
  • Previous data-labeling or model-evaluation work is helpful but not required.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Advanced Mathematics LLM Evaluation Expert

Join OpenTrain AI to design, solve, and evaluate challenging mathematics problems that probe LLM reasoning limits; this remote, part-time contractor role expects 20+ hours/week, advanced graduate-level math knowledge, Python ability, and English fluency.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026

Insurance LLM Evaluation SME (US, Remote)

Join OpenTrain as an Insurance LLM Evaluation SME to design and score underwriting, claims, and risk-assessment evaluation tasks for LLMs. Remote (U.S. only), $60–$80/hr, 35 hours/week, contractor/part-time.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Expert level
Hourly · $60–$80/hr

Posted Jul 10, 2026

LLM Agent Evaluation Scenario Writer

Design structured evaluation scenarios and gold-standard behaviors for LLM-based agents in a remote, part-time contractor role (20+ hrs/week). Pay $18–$24/hr; requires QA-style thinking, basic Python/JavaScript, and strong written English.

Generative AI & RLHF
Text
Remote · Worldwide
Part-time · Flexible
Intermediate level
Hourly · $18–$24/hr

Posted Jan 13, 2026