Skip to content
OpenTrain AIFor AI Companies

AI Evaluation Analyst — LLM Conversation & Rubric Authoring

Create multi-turn conversations, rubrics, and evaluation assets for frontier LLMs while working remotely as a contractor 20+ hours/week. Rapid onboarding and clear specs; paid on a per-task/hour basis at $20–$30/hr.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $20–$30/hr

$20–$30/hr

Compensation

Worldwide

Eligibility

Entry

Experience

Jul 16, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for building careers in AI training and data labeling. We help freelancers find and grow durable, remote work teaching AI by consolidating projects, tracking contributions, and building a single professional portfolio.

OpenTrain AI is the hiring and contracting organization for this role; we provide onboarding, project specifications, and quality guidance so contributors can focus on producing high-quality evaluation data.

  • Free to join and designed for flexible, remote work.
  • Work that directly shapes how state-of-the-art AI models behave.

About AI training work

AI training (also called data labeling or human feedback work) is the human side of building modern models: people prepare, test, and rate examples that models learn from. This role sits squarely in evaluation and RLHF workflows, generating the labeled content that improves model behavior.

Projects in this space are often part-time, remote, and accessible to contributors with strong domain skills such as writing, critical analysis, or model evaluation experience.

  • Work is remote and often flexible; many contributors choose hours that fit their schedule.
  • No single career path is required — attention to detail and clear writing are central.

The role

As an AI Evaluation Analyst you will author task-based multi-turn conversations, create aligned binary rubrics, test drafts against frontier LLMs, and deliver evaluation assets (transcripts, target behaviors, rubrics, supporting evidence). You will follow detailed specifications, calibrate as guidelines change, and validate outputs with team leads and QC.

  • Employment: Contractor, part-time.
  • Time: 20+ hours per week expected.
  • Languages: Native-level written English required.
  • Data & label types: Text data; TEXT_GENERATION, EVALUATION_RATING, RLHF.

What you'll do

This is writing- and analysis-heavy evaluation work. You will produce high-quality, specification-faithful artifacts that can be used to measure and train LLM behavior.

  • Author detailed, task-based multi-turn conversations for model evaluation.
  • Design and apply binary rubrics aligned to project goals and target behaviors.
  • Test and refine conversation drafts against frontier LLM outputs.
  • Deliver transcripts, rubric decisions, and supporting evidence for each item.
  • Maintain fidelity to changing specs and validate outputs with leads and QC.

Requirements

Candidates must meet the core requirements below; relevant experience is preferred but entry-level applicants with demonstrable writing and analytical skill are welcome.

  • Native-level written English with exceptional clarity, structure, and attention to detail.
  • Ability to interpret and follow highly detailed specifications independently.
  • Ability to write detailed multi-turn conversations and task-based rubrics.
  • Working knowledge of frontier LLM behaviors and common model failure patterns.
  • Experience in data annotation, RLHF, SFT, evaluation, or prompt engineering is preferred.
  • Helpful: research, editorial, technical writing, or QA experience; experience authoring evaluation items or analyzing model outputs.

Compensation & schedule

Compensation is output-focused and tied to meeting project specifications. This posting lists pay information as pay-per-hour with a range and expected output rates.

Onboarding and first tasks are available quickly; contributors typically begin work within 24–48 hours after onboarding.

  • Pay type: PAY_PER_HOUR (output-based expectations also apply).
  • Hourly range: $20–$30 USD per hour (projects may use minimum submission requirements).
  • Start: Rapid onboarding; first tasks expected within 24–48 hours.
  • Work style: Contractor, remote, consistent independent execution required.

Who should apply & how this works

Apply if you enjoy structured writing and evaluation, can work independently to spec, and have a strong eye for language and model behavior. This role is ideal for freelancers who want steady, remote, writing-heavy evaluation work in a cutting-edge AI setting.

OpenTrain contributors work on clearly specified tasks, validate outputs with leads and QC, and grow their portfolio through repeat projects and calibrated feedback.

  • Best fit: writers, editors, researchers, QA analysts, or anyone with experience in annotation or RLHF workflows.
  • Commitment: 20+ hours/week; contractor arrangement with output expectations.
  • Onboarding: Quick start with documented specs and QC steps.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

LLM Agent Evaluation Scenario Writer

Design structured evaluation scenarios and gold-standard behaviors for LLM-based agents in a remote, part-time contractor role (20+ hrs/week). Pay $18–$24/hr; requires QA-style thinking, basic Python/JavaScript, and strong written English.

Generative AI & RLHF
Text
Remote · Worldwide
Part-time · Flexible
Intermediate level
Hourly · $18–$24/hr

Posted Jan 13, 2026

Insurance LLM Evaluation SME (US, Remote)

Join OpenTrain as an Insurance LLM Evaluation SME to design and score underwriting, claims, and risk-assessment evaluation tasks for LLMs. Remote (U.S. only), $60–$80/hr, 35 hours/week, contractor/part-time.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Expert level
Hourly · $60–$80/hr

Posted Jul 10, 2026

English LLM Evaluation Generalist

Join OpenTrain to evaluate large language model outputs, create challenging prompts, and deliver recorded verbal feedback; remote, contract role (20+ hrs/week) paying $20–$30/hr. Entry-level friendly for strong American English speakers with LLM experience.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $20–$30/hr

Posted Jul 15, 2026