Test conversational AI models and agents for jailbreaks, prompt injections, bias exploitation, and other safety risks. This expert-level remote contract pays $48 to $62 per hour for 20+ hours weekly.
Generative AI & RLHF
100% Remote Hourly · $48–$62/hr
$48–$62/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Jul 30, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. Create a free account to build your AI training profile and apply in minutes.
OpenTrain helps specialists turn project-based AI work into a durable freelance career by making it easier to discover opportunities, track applications, and showcase relevant experience.
About AI Safety Training
AI training is the human side of building artificial intelligence. Specialists evaluate model behavior, identify failures, and provide structured feedback that helps make AI systems more reliable and safer.
As a red teaming specialist, you will work on the cutting edge of this field by challenging conversational models and agents with realistic adversarial scenarios. Your findings can help improve safety coverage across modern AI systems.
The Role
OpenTrain is hiring an expert AI Safety Red Teaming Specialist to probe conversational AI models and agents for vulnerabilities. This text-based work focuses on jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
You will surface weaknesses, classify model failures, flag systemic risks, and produce reproducible reports, datasets, and attack cases. Testing may involve sensitive topics such as bias, misinformation, and harmful behaviors, with structured guidelines and playbooks supporting consistent evaluation.
Expert-level AI safety red teaming contract
Remote work available worldwide
Part-time contractor engagement
20+ hours per week
Pay of $48 to $62 per hour
Work is text-based and focused on conversational AI
What You'll Do
You will apply adversarial thinking and structured evaluation methods to identify where conversational AI models and agents fail. Your work will help produce actionable safety insights for improving model behavior and evaluation coverage.
Red team conversational AI models and agents with adversarial prompts and multi-turn attacks
Annotate failures, classify vulnerabilities, and flag systemic risks in model behavior
Follow taxonomies, benchmarks, and playbooks to keep evaluations consistent and reproducible
Produce reports, datasets, and attack cases that can be used to improve safety coverage
Probe jailbreaks, prompt injections, misuse cases, and related model weaknesses
Required Qualifications
This role is designed for experienced practitioners who can investigate complex model behavior and communicate risk clearly. You must be comfortable working systematically with sensitive safety scenarios and documenting findings for both technical and non-technical audiences.
Prior experience with AI red teaming or adversarial machine learning
Experience in cybersecurity or socio-technical probing
Ability to identify jailbreaks, prompt injections, misuse cases, and related weaknesses
Experience using taxonomies, benchmarks, or structured evaluation playbooks
Strong structured thinking and clear written communication
Native fluency in English and Norwegian
Helpful Experience
The following backgrounds can strengthen your application and may help you contribute across a wider range of safety evaluations.
Experience with jailbreak datasets or prompt injection
Experience with RLHF or DPO attacks
Experience with model extraction
Background in penetration testing, exploit development, or reverse engineering
Experience probing harassment, disinformation, or other socio-technical risk scenarios
Why Join AI Training Work
AI training and data labeling offer flexible, remote ways to contribute to the development of advanced technology. Contributors with specialized expertise can work directly on challenging evaluation problems while choosing projects that fit their skills and availability.
Through OpenTrain, you can build a profile around your AI training experience and grow your career in a rapidly expanding industry where human judgment remains essential to improving AI.
Work remotely from anywhere with an internet connection
Choose flexible work that fits around other commitments
Apply specialized cybersecurity and AI safety expertise
Help shape how conversational AI systems behave
How to Apply
Create a free OpenTrain account, build your profile around your AI safety and red teaming experience, and apply through OpenTrain. Highlight your work with adversarial testing, structured evaluations, and technical risk communication.
Challenge frontier AI systems with adversarial prompts, uncover safety weaknesses, and document model behavior across high-risk topics. This expert contractor role pays $70 to $84 per hour for 20+ hours weekly.
Lead quality assurance for AI red-teaming and safety evaluation projects, reviewing adversarial prompts, risk analyses, and contributor work. This remote U.S. contract role offers up to $100 per hour and requires 20+ hours weekly.
Probe conversational AI for jailbreaks, prompt injections, bias exploitation, and manipulation as an expert red team contractor. Work worldwide for $48 to $62 per hour, 20+ hours weekly, using English and Danish.