Skip to content
OpenTrain AIFor AI Companies

LLM Evaluation Specialist

Evaluate large language models and help improve their performance in a remote, part-time freelance role. Open to students and graduates in any academic field, with pay up to $40 per hour.

Apply now
OpenTrain AI

Generative AI & RLHF

Remote Hourly · $40/hr

$40/hr

Compensation

1 country

Eligibility

Entry

Experience

Sep 30, 2026

Posted

Open to applicants in

United States

The work

As an LLM Evaluation Specialist, you will test large language models and assess how well they perform. You will use careful reasoning, follow detailed instructions, and review model behavior accurately while working independently in a mainly asynchronous setting.

  • Test large language models as part of AI training projects.
  • Use general academic knowledge and reasoning to evaluate model performance.
  • Follow detailed instructions while completing evaluation tasks.
  • Review model behavior carefully and provide work that can support better system performance.
  • Work independently in a remote, primarily asynchronous environment.

What it pays and takes

This is a part-time freelance contractor role for short-term AI training projects. The work is suitable for students and graduates from any academic discipline, and no specific professional or AI training background is required.

  • Pay: Up to $40 per hour.
  • Work type: Remote, part-time freelance contract work.
  • Location: United States.
  • Language: English.
  • Hours: The structured role details list 20+ hours per week; the role description says there is no minimum weekly commitment.
  • Education: Current enrollment in an associate, bachelor's, or master's program, or completion of a degree.
  • Requirements: Strong reasoning, exceptional attention to detail, consistent accuracy, and comfort working independently.
  • Training: Openness to brief training and learning new evaluation workflows.

How it works

Apply on OpenTrain with your resume and then complete the application on the hiring site.

About AI training work

AI training is the human work behind systems such as language models, including testing responses and rating model performance. OpenTrain hires and contracts contributors for this work, and people with strong judgment are paid to help make AI systems more useful and reliable.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Political Science LLM Evaluation Expert

Use political science expertise to create evaluation prompts, review LLM responses, test difficult cases, and give evidence-backed feedback. This eight-week contractor assignment requires 40 hours per week and four hours of Pacific Time overlap.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 15, 2026

Vietnamese LLM Prompt Evaluation Specialist

Create and assess Vietnamese prompts that test retrieval, reasoning, calculations, and instruction following. Review AI responses against clear quality standards in a remote, flexible contract role.

Generative AI & RLHF
Text
Remote · Worldwide
Vietnamese, English
Part-time · Flexible
Entry level

Posted Sep 22, 2026

LLM Conversation Evaluation Engineering Manager

Review multi-turn LLM conversations and tool-use scenarios, improve model responses, and explain clear evaluation decisions. This remote contractor role requires English fluency, independent work, and PST overlap.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Sep 16, 2026