Skip to content
OpenTrain AIFor AI Companies

AI Safety Practitioner – Model Evaluation

Evaluate frontier AI responses for safety, accuracy, policy compliance, and quality across high-risk topics while helping improve model alignment through structured RLHF, SFT, and safety benchmarking.

OpenTrain AI

Generative AI & RLHF

Remote Hourly · $60–$70/hr

$60–$70/hr

Compensation

33 countries

Eligibility

Expert

Experience

Jul 22, 2026

Posted

Open to applicants in

United States Denmark Estonia Finland Ireland Latvia Lithuania Norway Sweden Austria Belgium France Germany Netherlands Switzerland United Kingdom Albania Bosnia & Herzegovina Croatia Greece Italy Malta Portugal Serbia Slovenia Spain Bulgaria Czechia Hungary Moldova Poland Romania Slovakia

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and apply to opportunities in minutes.

Creating an OpenTrain account is free, and the platform is designed to help contributors grow in a fast-moving field at the center of how modern AI systems are built.

About AI Safety and Model Evaluation

AI training is the human side of building artificial intelligence. Expert contributors review model outputs, identify failures, and provide structured feedback so AI systems become more useful, accurate, aligned, and safe.

This role focuses on generative AI evaluation and safety work, including reviewing sensitive outputs and applying rubrics used in reinforcement learning from human feedback (RLHF), supervised fine-tuning (SFT), and AI safety benchmarking.

  • Work on cutting-edge AI systems and model behavior
  • Use analytical judgment to assess nuanced and policy-sensitive content
  • Flexible, part-time contracting work requiring 20 or more hours per week

The AI Safety Practitioner Role

OpenTrain AI is recruiting an expert AI Safety Practitioner to evaluate frontier AI model responses for safety, factual accuracy, policy compliance, and overall quality. You will assess ambiguous and high-risk content across areas such as misinformation, political persuasion, self-harm, violence, cyber, and biosecurity.

Your evaluations will help identify unsafe outputs, hallucinations, reasoning failures, and policy violations. You will then provide structured feedback that supports model alignment and improves safety performance.

  • Expert-level role for applicants with substantial professional experience
  • Part-time contractor engagement
  • Expected commitment of 20+ hours per week
  • Pay range: $60–$70 USD per hour

What You'll Do

You will apply established safety policies and quality standards to AI-generated responses, make careful judgments in ambiguous cases, and document findings clearly. The work includes collaboration with AI researchers and safety teams on ongoing evaluation initiatives.

  • Evaluate AI-generated responses against safety policies and quality standards
  • Review ambiguous and high-risk content across sensitive subject areas
  • Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking
  • Identify unsafe outputs, hallucinations, reasoning failures, and policy violations
  • Provide structured feedback to support model alignment and safety improvement
  • Collaborate with AI researchers and safety teams on ongoing evaluation work

Required Qualifications

Applicants should bring advanced professional experience assessing nuanced, policy-sensitive, or high-risk content. A bachelor's degree or higher is required in a related discipline, along with excellent written English, critical thinking, and analytical reasoning.

  • Bachelor’s degree or higher in journalism, communications, psychology, sociology, public policy, law, biology, chemistry, computer science, or a related discipline
  • 5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related field
  • Experience evaluating AI-generated responses for safety and quality
  • Experience evaluating nuanced, policy-sensitive, or high-risk content
  • Strong analytical judgment in ambiguous, policy-sensitive scenarios
  • Excellent written English, critical thinking, and analytical reasoning
  • Familiarity with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluation preferred
  • Professional background in AI Safety, Trust & Safety, policy, research, or security

Eligible Locations and Language

This opportunity is intended for contractors in the eligible countries listed below. English is the required working language.

  • Language: English
  • Eligible countries: United States, Denmark, Estonia, Finland, Ireland, Latvia, Lithuania, Norway, Sweden, Austria, Belgium, France, Germany, Netherlands, Switzerland, United Kingdom
  • Eligible countries: Albania, Bosnia and Herzegovina, Croatia, Greece, Italy, Malta, Portugal, Serbia, Slovenia, Spain, Bulgaria, Czechia, Hungary, Moldova, Poland, Romania, Slovakia

Build Your AI Training Career with OpenTrain

AI safety evaluation is part of a rapidly growing industry where human expertise directly shapes how advanced models behave. Through OpenTrain, you can create a profile, discover AI training opportunities, and build experience in work that sits at the forefront of artificial intelligence.

  • Create a free OpenTrain account
  • Build a profile highlighting your safety, policy, research, or technical experience
  • Apply to relevant AI training and evaluation opportunities

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Bilingual LLM Safety Evaluator (Hebrew & English)

Join OpenTrain AI as a remote, part-time contractor reviewing and red-teaming LLM outputs in Hebrew and English to find safety failures and produce labeled evaluation data. $26–$38/hr, 20+ hours/week; your feedback will directly shape model safety.

Generative AI & RLHF
Text
Remote · Worldwide
Part-time · Flexible
Intermediate level
Hourly · $26–$38/hr

Posted Apr 3, 2026

Nuclear & Radiological Security Expert (RLHF Evaluation)

Join OpenTrain AI to define safety standards, escalation protocols, and abstraction frameworks for nuclear and radiological risk evaluation; remote, 20+ hrs/week, $50–$90/hr. Help shape how advanced models handle sensitive nuclear security information.

Generative AI & RLHF
Document
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $50–$90/hr

Posted Jun 30, 2026

Bilingual LLM Safety Evaluator (French/English)

Join OpenTrain as a remote contractor to evaluate and red-team LLM outputs in French and English, focusing on safety, policy alignment, and adversarial case curation. This part-time role (20+ hrs/week) pays $24–$36/hr (typical $30/hr) and requires hands-on LLM red-teaming experience.

Generative AI & RLHF
Text
Remote · Worldwide
Part-time · Flexible
Intermediate level
Hourly · $24–$36/hr

Posted Apr 3, 2026