Skip to content
OpenTrain AIFor AI Companies

Insurance LLM Evaluation SME (US, Remote)

Join OpenTrain as an Insurance LLM Evaluation SME to design and score underwriting, claims, and risk-assessment evaluation tasks for LLMs. Remote (U.S. only), $60–$80/hr, 35 hours/week, contractor/part-time.

OpenTrain AI

Generative AI & RLHF

Remote Hourly · $60–$80/hr

$60–$80/hr

Compensation

1 country

Eligibility

Expert

Experience

Jul 10, 2026

Posted

Open to applicants in

United States

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people start and grow careers teaching AI by bringing specialized projects and evaluation work into one place and supporting contractors as they develop a long-term portfolio.

  • OpenTrain is the hiring and contracting organization for this role.
  • Creating an OpenTrain account is free and is how you apply and manage your work.

Why AI training in insurance matters

AI training (data labeling and human evaluation) is how modern AI systems learn to reason and act. In insurance, expert human feedback helps models handle underwriting tradeoffs, assess claims fairly, and reason about risk — work that shapes how AI supports day-to-day insurance decisions.

This role puts experienced insurance professionals on the front lines of model behavior: you’ll convert real-world judgment into evaluation tasks and scoring guidance that teach LLMs to reason like a subject-matter expert.

  • Flexible, remote work that directly impacts model safety and reliability.
  • Work that rewards deep domain judgment and clear, evidence-based feedback.

The role — what an Insurance LLM Evaluation SME does

As an Insurance LLM Evaluation SME you will design domain-focused evaluation tasks, author grounded solutions, and score LLM outputs against structured rubrics. You’ll be responsible for clear, written feedback on correctness, judgment, and reasoning quality, plus refining scoring guidelines to keep data consistent across subject-matter experts.

  • Design scenarios and prompts that probe underwriting, claims, and risk-assessment reasoning.
  • Write accurate, practice-grounded answers and model solutions for evaluation use.
  • Review and score model-generated text using rubrics; provide concise, actionable written feedback.
  • Help refine evaluation guidelines and maintain consistency across SMEs.

What you'll do day-to-day

Work with fellow SMEs and the project team to create and maintain high-quality evaluation sets. Produce clear written solutions and scoring notes that capture real-world insurance judgment. Regularly score LLM responses, document edge cases, and suggest rubric updates when models surface new failure modes.

  • Create and iterate evaluation tasks focused on underwriting, claims, actuarial, or risk scenarios.
  • Score text outputs (evaluation rating and text-generation tasks) against established rubrics.
  • Draft and edit scoring guidelines and example solutions used by other evaluators.
  • Communicate findings and consistency issues to improve rubric reliability.

Requirements

We require deep, hands-on insurance domain expertise and demonstrable experience applying professional judgment. Candidates must be able to produce clear, written solutions and evaluate LLM outputs against structured scoring criteria.

  • Deep professional experience in insurance (underwriting, claims, actuarial work, or risk management).
  • Hands-on experience evaluating AI or LLM outputs against rubrics or structured scoring frameworks.
  • Strong written communication and ability to explain reasoning clearly and concisely.
  • Availability to work reliably 35 hours per week, on weekdays (U.S. time zones).
  • Must be based in the United States and fluent in English.

Compensation, commitment, and logistics

This is a contract, part-time position with a fixed hourly range and structured evaluation responsibilities. Pay and scheduling are set so subject-matter experts can focus on high-quality, consistent review work.

  • Pay: $60–$80 per hour (USD).
  • Commitment: 35 hours per week, weekdays required.
  • Employment type: Contractor, part-time.
  • Work location: Remote (U.S. only).
  • Data & label types: Text data; evaluation rating and text-generation tasks.

How to apply

If you meet the requirements and want to shape how insurance-focused LLMs reason, apply through OpenTrain. Create or sign in to your OpenTrain account, complete your profile, and submit your application so the hiring team can review your background and domain experience.

  • Prepare examples of relevant insurance work and any experience evaluating model outputs or using rubrics.
  • OpenTrain will contact qualified candidates with next steps for onboarding and sample evaluation tasks.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

LLM Safety Evaluator (Hebrew & English Required)

Join OpenTrain AI as a remote, part-time contractor reviewing and red-teaming LLM outputs in Hebrew and English to find safety failures and produce labeled evaluation data. $26–$38/hr, 20+ hours/week; your feedback will directly shape model safety.

Generative AI & RLHF
Text
Remote · Worldwide
Part-time · Flexible
Intermediate level
Hourly · $26–$38/hr

Posted Apr 3, 2026

Finance LLM Evaluation & Rubric Design Expert

Join OpenTrain AI to shape how finance-focused LLMs reason: design domain-realistic tasks, build scoring rubrics, and evaluate model outputs with detailed written feedback. This remote, US-based contract role expects 35 hours/week at $65–$90/hr.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Expert level
Hourly · $65–$90/hr

Posted Jul 13, 2026

Evaluation Scenario Writer - AI Agent Testing Specialist

Design structured evaluation scenarios and gold-standard behaviors for LLM-based agents in a remote, part-time contractor role (20+ hrs/week). Pay $18–$24/hr; requires QA-style thinking, basic Python/JavaScript, and strong written English.

Generative AI & RLHF
Text
Remote · Worldwide
Part-time · Flexible
Intermediate level
Hourly · $18–$24/hr

Posted Jan 13, 2026