Skip to content
OpenTrain AIFor AI Companies

Science & Technology LLM Evaluator

Use advanced science and technology expertise to design challenging prompts, evaluate LLM responses, and provide evidence-based feedback. This remote, part-time contractor role is open worldwide and requires 20+ hours per week.

OpenTrain AI

Generative AI & RLHF

100% Remote

Worldwide

Eligibility

Entry

Experience

Aug 6, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and grow in a rapidly expanding field. Creating an OpenTrain account is free.

About AI Training Work

AI training is the human side of building modern artificial intelligence. Specialists create examples, assess model responses, and identify errors so that large language models become more accurate, reliable, and useful. In this role, your scientific and technical judgment will directly contribute to improving advanced AI systems.

  • Work remotely with flexible, part-time contractor scheduling.
  • Apply real scientific and technical expertise to cutting-edge generative AI.
  • Help shape how AI systems reason about complex scientific and technological topics.

The Role

OpenTrain AI is seeking a Science & Technology LLM Evaluator to assess and improve large language models. You will create demanding prompts, review AI-generated responses, and provide objective, evidence-backed feedback across scientific and technical subjects.

The role is classified as entry level in the project information, while the stated requirements include a master's degree or higher and at least three years of professional, research, or teaching experience in a relevant field.

  • Employment type: Part-time contractor
  • Time requirement: 20+ hours per week
  • Location: Worldwide and fully remote
  • Primary language: English
  • Work type: Text-based prompt writing and response evaluation

What You'll Do

You will design challenging evaluation content and review model performance for factual accuracy, reasoning quality, completeness, and nuance. Your work will help uncover weaknesses in AI systems and generate high-quality training and benchmarking data.

  • Create advanced prompts covering physics, chemistry, biology, engineering, AI, computer science, and emerging technologies.
  • Evaluate AI-generated responses for factual accuracy, reasoning quality, completeness, and nuance.
  • Identify hallucinations, logical inconsistencies, outdated information, and edge cases.
  • Develop benchmark datasets and adversarial test cases.
  • Provide evidence-based feedback supported by reliable references.
  • Collaborate with AI researchers to improve model performance.
  • Maintain high annotation quality and thorough documentation.

Requirements

This opportunity is designed for candidates with advanced academic or practical expertise in science and technology. You must be able to work independently, apply sound analytical judgment, and explain evaluations clearly in written English.

  • Master's degree or higher in a science or technology discipline, such as physics, chemistry, biology, engineering, or computer science.
  • At least three years of professional, research, or teaching experience in your field.
  • Excellent written English.
  • Strong analytical and critical-thinking skills.
  • Experience evaluating AI model outputs for factual accuracy and reasoning.
  • Ability to design challenging prompts that test model limits.
  • Strong attention to detail.
  • Ability to provide objective, evidence-backed reviews.
  • Commitment to high-quality annotation and documentation.

Helpful Background

The following experience is helpful but is listed separately from the required qualifications.

  • Experience with large language models or generative AI.
  • Prompt engineering experience.
  • Previous AI evaluation work.
  • Published research.
  • Industry recognition.
  • Teaching experience.

Why This Work Matters

Every major AI system depends on people who can prepare, review, and improve the examples used for training. By evaluating scientific and technical reasoning, you will help make AI outputs more trustworthy in fields where accuracy and nuance matter.

  • Contribute to the development of state-of-the-art AI models.
  • Work from anywhere with an internet-connected computer.
  • Build experience in a fast-growing AI training industry.
  • Fit flexible project work around other professional or personal commitments.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Physics LLM Evaluation Expert, STEM Problem Design

Join OpenTrain to design and solve advanced physics problems and build evaluation benchmarks that fine-tune large language models. This remote, part-time contractor role (20+ hrs/week) suits PhD-level physicists or equivalent with strong symbolic and multi-step reasoning skills.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Expert level

Posted Jul 17, 2026

Physics LLM Evaluation Expert

Join OpenTrain to design and solve advanced physics problems that probe LLM reasoning and symbolic skills; remote, part-time contractor work (20+ hrs/week) for candidates with graduate-level physics experience and strong written English.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Expert level

Posted Jul 17, 2026

Physics LLM Evaluation Expert

Create and solve demanding physics problems that help evaluate and improve large language models. This fully remote, part-time contractor role is open worldwide and offers advanced physics students the chance to shape cutting-edge AI systems.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 7, 2026