Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Teaming Specialist

Test conversational AI models and agents for jailbreaks, prompt injections, bias exploitation, and other safety risks. This expert-level remote contract pays $48 to $62 per hour for 20+ hours weekly.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $48–$62/hr

$48–$62/hr

Compensation

Worldwide

Eligibility

Expert

Experience

Jul 30, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. Create a free account to build your AI training profile and apply in minutes.

OpenTrain helps specialists turn project-based AI work into a durable freelance career by making it easier to discover opportunities, track applications, and showcase relevant experience.

About AI Safety Training

AI training is the human side of building artificial intelligence. Specialists evaluate model behavior, identify failures, and provide structured feedback that helps make AI systems more reliable and safer.

As a red teaming specialist, you will work on the cutting edge of this field by challenging conversational models and agents with realistic adversarial scenarios. Your findings can help improve safety coverage across modern AI systems.

The Role

OpenTrain is hiring an expert AI Safety Red Teaming Specialist to probe conversational AI models and agents for vulnerabilities. This text-based work focuses on jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.

You will surface weaknesses, classify model failures, flag systemic risks, and produce reproducible reports, datasets, and attack cases. Testing may involve sensitive topics such as bias, misinformation, and harmful behaviors, with structured guidelines and playbooks supporting consistent evaluation.

  • Expert-level AI safety red teaming contract
  • Remote work available worldwide
  • Part-time contractor engagement
  • 20+ hours per week
  • Pay of $48 to $62 per hour
  • Work is text-based and focused on conversational AI

What You'll Do

You will apply adversarial thinking and structured evaluation methods to identify where conversational AI models and agents fail. Your work will help produce actionable safety insights for improving model behavior and evaluation coverage.

  • Red team conversational AI models and agents with adversarial prompts and multi-turn attacks
  • Annotate failures, classify vulnerabilities, and flag systemic risks in model behavior
  • Follow taxonomies, benchmarks, and playbooks to keep evaluations consistent and reproducible
  • Produce reports, datasets, and attack cases that can be used to improve safety coverage
  • Probe jailbreaks, prompt injections, misuse cases, and related model weaknesses

Required Qualifications

This role is designed for experienced practitioners who can investigate complex model behavior and communicate risk clearly. You must be comfortable working systematically with sensitive safety scenarios and documenting findings for both technical and non-technical audiences.

  • Prior experience with AI red teaming or adversarial machine learning
  • Experience in cybersecurity or socio-technical probing
  • Ability to identify jailbreaks, prompt injections, misuse cases, and related weaknesses
  • Experience using taxonomies, benchmarks, or structured evaluation playbooks
  • Strong structured thinking and clear written communication
  • Native fluency in English and Norwegian

Helpful Experience

The following backgrounds can strengthen your application and may help you contribute across a wider range of safety evaluations.

  • Experience with jailbreak datasets or prompt injection
  • Experience with RLHF or DPO attacks
  • Experience with model extraction
  • Background in penetration testing, exploit development, or reverse engineering
  • Experience probing harassment, disinformation, or other socio-technical risk scenarios

Why Join AI Training Work

AI training and data labeling offer flexible, remote ways to contribute to the development of advanced technology. Contributors with specialized expertise can work directly on challenging evaluation problems while choosing projects that fit their skills and availability.

Through OpenTrain, you can build a profile around your AI training experience and grow your career in a rapidly expanding industry where human judgment remains essential to improving AI.

  • Work remotely from anywhere with an internet connection
  • Choose flexible work that fits around other commitments
  • Apply specialized cybersecurity and AI safety expertise
  • Help shape how conversational AI systems behave

How to Apply

Create a free OpenTrain account, build your profile around your AI safety and red teaming experience, and apply through OpenTrain. Highlight your work with adversarial testing, structured evaluations, and technical risk communication.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

AI Safety Red Teamer

Challenge frontier AI systems with adversarial prompts, uncover safety weaknesses, and document model behavior across high-risk topics. This expert contractor role pays $70 to $84 per hour for 20+ hours weekly.

Generative AI & RLHF
Text
Remote · United States, Denmark, Estonia +30 more
English
Part-time · Flexible
Expert level
Hourly · $70–$84/hr

Posted Jul 17, 2026

Red-Teaming Quality Assurance Lead

Lead quality assurance for AI red-teaming and safety evaluation projects, reviewing adversarial prompts, risk analyses, and contributor work. This remote U.S. contract role offers up to $100 per hour and requires 20+ hours weekly.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Expert level
Hourly · $100/hr

Posted Jul 9, 2026

AI Safety Red Team Expert

Probe conversational AI for jailbreaks, prompt injections, bias exploitation, and manipulation as an expert red team contractor. Work worldwide for $48 to $62 per hour, 20+ hours weekly, using English and Danish.

Generative AI & RLHF
Text
Remote · Worldwide
English, Danish
Part-time · Flexible
Expert level
Hourly · $48–$62/hr

Posted Jul 30, 2026