Evaluate frontier AI responses for safety, accuracy, policy compliance, and quality across high-risk topics while helping improve model alignment through structured RLHF, SFT, and safety benchmarking.
Generative AI & RLHF
Remote Hourly · $60–$70/hr
$60–$70/hr
Compensation
33 countries
Eligibility
Expert
Experience
Jul 22, 2026
Posted
Open to applicants in
United States Denmark Estonia Finland Ireland Latvia Lithuania Norway Sweden Austria Belgium France Germany Netherlands Switzerland United Kingdom Albania Bosnia & Herzegovina Croatia Greece Italy Malta Portugal Serbia Slovenia Spain Bulgaria Czechia Hungary Moldova Poland Romania Slovakia
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and apply to opportunities in minutes.
Creating an OpenTrain account is free, and the platform is designed to help contributors grow in a fast-moving field at the center of how modern AI systems are built.
About AI Safety and Model Evaluation
AI training is the human side of building artificial intelligence. Expert contributors review model outputs, identify failures, and provide structured feedback so AI systems become more useful, accurate, aligned, and safe.
This role focuses on generative AI evaluation and safety work, including reviewing sensitive outputs and applying rubrics used in reinforcement learning from human feedback (RLHF), supervised fine-tuning (SFT), and AI safety benchmarking.
Work on cutting-edge AI systems and model behavior
Use analytical judgment to assess nuanced and policy-sensitive content
Flexible, part-time contracting work requiring 20 or more hours per week
The AI Safety Practitioner Role
OpenTrain AI is recruiting an expert AI Safety Practitioner to evaluate frontier AI model responses for safety, factual accuracy, policy compliance, and overall quality. You will assess ambiguous and high-risk content across areas such as misinformation, political persuasion, self-harm, violence, cyber, and biosecurity.
Your evaluations will help identify unsafe outputs, hallucinations, reasoning failures, and policy violations. You will then provide structured feedback that supports model alignment and improves safety performance.
Expert-level role for applicants with substantial professional experience
Part-time contractor engagement
Expected commitment of 20+ hours per week
Pay range: $60–$70 USD per hour
What You'll Do
You will apply established safety policies and quality standards to AI-generated responses, make careful judgments in ambiguous cases, and document findings clearly. The work includes collaboration with AI researchers and safety teams on ongoing evaluation initiatives.
Evaluate AI-generated responses against safety policies and quality standards
Review ambiguous and high-risk content across sensitive subject areas
Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking
Identify unsafe outputs, hallucinations, reasoning failures, and policy violations
Provide structured feedback to support model alignment and safety improvement
Collaborate with AI researchers and safety teams on ongoing evaluation work
Required Qualifications
Applicants should bring advanced professional experience assessing nuanced, policy-sensitive, or high-risk content. A bachelor's degree or higher is required in a related discipline, along with excellent written English, critical thinking, and analytical reasoning.
Bachelor’s degree or higher in journalism, communications, psychology, sociology, public policy, law, biology, chemistry, computer science, or a related discipline
5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related field
Experience evaluating AI-generated responses for safety and quality
Experience evaluating nuanced, policy-sensitive, or high-risk content
Strong analytical judgment in ambiguous, policy-sensitive scenarios
Excellent written English, critical thinking, and analytical reasoning
Familiarity with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluation preferred
Professional background in AI Safety, Trust & Safety, policy, research, or security
Eligible Locations and Language
This opportunity is intended for contractors in the eligible countries listed below. English is the required working language.
Language: English
Eligible countries: United States, Denmark, Estonia, Finland, Ireland, Latvia, Lithuania, Norway, Sweden, Austria, Belgium, France, Germany, Netherlands, Switzerland, United Kingdom
Eligible countries: Albania, Bosnia and Herzegovina, Croatia, Greece, Italy, Malta, Portugal, Serbia, Slovenia, Spain, Bulgaria, Czechia, Hungary, Moldova, Poland, Romania, Slovakia
Build Your AI Training Career with OpenTrain
AI safety evaluation is part of a rapidly growing industry where human expertise directly shapes how advanced models behave. Through OpenTrain, you can create a profile, discover AI training opportunities, and build experience in work that sits at the forefront of artificial intelligence.
Create a free OpenTrain account
Build a profile highlighting your safety, policy, research, or technical experience
Apply to relevant AI training and evaluation opportunities
Join OpenTrain AI as a remote, part-time contractor reviewing and red-teaming LLM outputs in Hebrew and English to find safety failures and produce labeled evaluation data. $26–$38/hr, 20+ hours/week; your feedback will directly shape model safety.
Join OpenTrain AI to define safety standards, escalation protocols, and abstraction frameworks for nuclear and radiological risk evaluation; remote, 20+ hrs/week, $50–$90/hr. Help shape how advanced models handle sensitive nuclear security information.
Join OpenTrain as a remote contractor to evaluate and red-team LLM outputs in French and English, focusing on safety, policy alignment, and adversarial case curation. This part-time role (20+ hrs/week) pays $24–$36/hr (typical $30/hr) and requires hands-on LLM red-teaming experience.