Use deep insurance expertise to evaluate and improve generative AI outputs across underwriting, claims, and risk assessment. This remote contract pays $60–$80 per hour and requires prior LLM evaluation experience.
Generative AI & RLHF
Remote Hourly · $60–$80/hr
$60–$80/hr
Compensation
1 country
Eligibility
Expert
Experience
Jul 10, 2026
Posted
Open to applicants in
United States
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping experts discover projects, build a professional profile, and apply to cutting-edge work.
Free account creation and a streamlined application process
Opportunities to build a lasting career in AI training and data labeling
About AI Training Work
AI training is the human side of building modern artificial intelligence. Subject-matter experts review model outputs, create challenging examples, and provide structured feedback that helps AI systems reason more accurately in specialized fields such as insurance.
Help shape how advanced AI models evaluate insurance scenarios
Apply professional judgment to real-world model responses
Work remotely in a fast-growing technology field
The Role
OpenTrain AI is seeking an Insurance AI Model Evaluator to support generative AI model improvement. You will apply expert knowledge from underwriting, claims, actuarial work, or risk management to design domain-specific tasks, assess model reasoning, and deliver precise feedback.
This is an expert-level contractor and part-time engagement for professionals based in the United States. The structured project requirement is 20+ hours per week, with reliable availability of at least 35 weekday hours expected for the role.
Pay: $60–$80 per hour
Engagement: Contractor, part-time
Location: United States
Primary language: English
Data type: Text
Work type: Evaluation and rating of AI outputs
What You'll Do
You will help research and engineering teams identify and close gaps in how AI models handle insurance reasoning. Your work will combine task design, expert review, rubric development, and written communication.
Guide teams on knowledge gaps involving underwriting, claims, and risk-assessment reasoning.
Write accurate, well-reasoned solutions grounded in real underwriting or claims practice.
Evaluate AI model outputs using structured rubrics.
Provide clear written feedback on correctness, judgment, and reasoning quality.
Develop and refine evaluation guidelines and insurance-specific scoring rubrics.
Collaborate with other subject-matter experts to maintain consistency and accuracy in training data.
Required Qualifications
This role requires substantial professional insurance experience and demonstrated familiarity with evaluating large language model or AI outputs. Candidates should be able to communicate nuanced insurance judgment clearly and apply structured scoring criteria consistently.
8+ years of dedicated professional insurance experience at a recognized, top-tier organization.
Background in underwriting, claims, actuarial work, risk management, or a related insurance function.
Mandatory hands-on experience evaluating LLM or AI model outputs against rubrics or structured scoring criteria.
Demonstrable career progression in insurance.
Strong verbal and written communication skills.
Strong problem-solving and interpersonal skills.
Reliable availability for at least 35 hours per week during weekdays.
Who Should Apply
This opportunity is designed for experienced insurance professionals who want to use their industry judgment to influence how generative AI systems perform. It may suit senior practitioners with a strong record of progression and direct experience reviewing AI outputs in a structured way.
Senior insurance specialists with deep practical expertise
Professionals experienced in underwriting, claims, actuarial, or risk functions
Experts comfortable writing detailed evaluations and reasoned solutions
Candidates located in the United States who can meet the weekday availability requirement
Use senior insurance expertise to create realistic enterprise scenarios, reference answers, and evaluation rubrics for AI systems. This remote contractor role pays $50–$60 per hour and requires 20+ hours per week.
Evaluate AI-generated responses, apply expert judgment, and provide feedback that improves model behavior. This flexible, worldwide contractor role offers 20+ hours per week and pays $140–$200 USD per hour.
Evaluate frontier AI responses for safety, accuracy, policy compliance, and quality across high-risk topics while helping improve model alignment through structured RLHF, SFT, and safety benchmarking.