Skip to content
OpenTrain AIFor AI Companies

LLM Evaluation Domain Reviewer

Review domain-specific LLM prompts and completed evaluations for accuracy, reasoning, completeness, and guideline compliance. This 8-week contractor assignment requires a master's degree, three years of expertise, and 40 hours per week.

OpenTrain AI

Generative AI & RLHF

100% Remote

Worldwide

Eligibility

Entry

Experience

Aug 4, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this opportunity. OpenTrain is the leading platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes.

Creating an OpenTrain account is free, and this role offers the opportunity to contribute directly to the quality of large language model evaluation data.

About AI Training Work

AI training is the human work behind modern artificial intelligence. Contributors write, review, and evaluate examples that help models produce more accurate, useful, and reliable responses.

In this role, your subject matter expertise will help assess model evaluation tasks across academic, research, and industry disciplines. This is remote contractor work in a rapidly growing field where human judgment directly shapes how AI systems behave.

The Role

OpenTrain is seeking an LLM Evaluation Domain Reviewer to support quality assurance for large language model evaluation projects. You will review domain-specific prompts, assess completed tasks, verify compliance with project guidelines, and provide actionable feedback to maintain high annotation standards.

The listing identifies this opportunity as entry level, while the role requires a master's degree or higher and at least three years of professional, research, teaching, or industry experience in your area of expertise.

  • Contractor assignment
  • Part-time employment classification
  • Eight-week project duration
  • English-language work
  • Worldwide opportunity

What You'll Do

You will apply careful analysis and subject matter knowledge to evaluate the quality and reliability of AI training data. Reviews should be consistent, evidence-based, and aligned with established project standards.

  • Review and validate domain-specific prompts within your area of expertise.
  • Evaluate completed tasks for factual accuracy, reasoning quality, completeness, and compliance with guidelines.
  • Identify factual inaccuracies, logical inconsistencies, hallucinations, outdated information, and low-quality annotations.
  • Check that prompts are challenging, relevant, and aligned with project objectives.
  • Provide clear, constructive, evidence-based feedback to contributors.
  • Maintain consistency across reviews by following established quality standards.
  • Escalate ambiguous or complex cases when necessary and document review findings.
  • Collaborate with project managers and AI teams to improve evaluation quality.

Requirements

Applicants should bring advanced education and meaningful experience in a relevant area of expertise. Prior experience reviewing, evaluating, editing, or quality-checking specialized content is helpful but preferred rather than required.

Strong written English, analytical judgment, and attention to detail are essential for identifying factual, logical, and contextual errors.

  • Master's degree or higher in any field.
  • At least three years of professional, research, teaching, or industry experience in your area of expertise.
  • Excellent written English communication skills.
  • Strong analytical skills and attention to detail.
  • Ability to identify factual, logical, and contextual errors.
  • Prior domain-specific content review or evaluation experience is helpful.

Schedule and Assignment Details

The project requires a 40-hour-per-week commitment with at least four hours of overlap during Pacific Standard Time. The structured opportunity is also listed as requiring 20 or more hours per week, so applicants should review the project schedule carefully before committing.

This is an eight-week contractor assignment. The assignment does not include medical leave or paid leave.

  • Project duration: 8 weeks
  • Project schedule: 40 hours per week
  • Required availability: at least 4 hours of PST overlap
  • Structured time requirement: 20+ hours per week
  • Assignment type: contractor
  • Medical leave and paid leave are not provided

How to Apply

Create a free OpenTrain account, build your profile around your education and professional expertise, and apply to the opportunity in minutes. Highlight experience with research, teaching, professional practice, or domain-specific content review so your qualifications are easy to assess.

If selected, you will use your subject matter knowledge to review prompts and model evaluation work that supports more accurate and dependable AI systems.

  • Create or update your free OpenTrain profile.
  • Showcase your master's degree or higher qualification.
  • Describe at least three years of relevant expertise.
  • Confirm your English communication skills and schedule availability.
  • Apply through OpenTrain.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Music Domain Reviewer for LLM Evaluations

Use your music expertise to review prompts and evaluate AI-generated content for accuracy, reasoning, and quality. This remote contract role offers 20+ hours per week through OpenTrain.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 6, 2026

Art Domain Reviewer for AI Evaluation

Use your art expertise to review prompts and evaluate AI training tasks across art history, visual arts, architecture, design, and museums. This 8-week contractor role offers 20+ hours weekly, with the role description specifying a 40-hour schedule.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 19, 2026

Legal LLM Evaluation Analyst

Use your legal reasoning, research, and writing skills to evaluate large language model outputs in a remote, one-month freelance project for contributors in India.

Generative AI & RLHF
Document
Remote · India
English
Part-time · Flexible
Entry level

Posted Aug 7, 2026