Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Team Expert

Test conversational AI for jailbreaks, prompt injections, bias exploitation, and multi-turn manipulation. This expert contract role offers $48-$62 per hour and requires native English and Danish fluency.

OpenTrain AI

Generative AI & RLHF

Remote Hourly · $48–$62/hr

$48–$62/hr

Compensation

30 countries

Eligibility

Expert

Experience

Jul 30, 2026

Posted

Open to applicants in

Austria Belgium Bulgaria
+27 more
  • Austria
  • Belgium
  • Bulgaria
  • Canada
  • Croatia
  • Cyprus
  • Czechia
  • Denmark
  • Estonia
  • Finland
  • France
  • Germany
  • Greece
  • Hungary
  • Ireland
  • Italy
  • Latvia
  • Lithuania
  • Luxembourg
  • Malta
  • Netherlands
  • Poland
  • Portugal
  • Romania
  • Slovakia
  • Slovenia
  • Spain
  • Sweden
  • United Kingdom
  • United States

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring contractors for specialized projects that help shape how advanced AI systems behave.

Create a free OpenTrain account to build a profile, showcase relevant experience, discover projects that match your skills, and apply in minutes.

About AI Safety Red Teaming

AI safety red teaming is a specialized form of AI training and evaluation. Human experts deliberately probe models with adversarial prompts and realistic misuse scenarios to uncover weaknesses before those systems are widely used.

This work contributes directly to safer generative AI by identifying failures, documenting risks, and producing examples that can improve model behavior.

The Role

OpenTrain is seeking an AI Safety Red Team Expert to test conversational AI models and agents. The work is text-based and focuses on finding jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.

You will use structured taxonomies, benchmarks, and playbooks to keep testing consistent across scenarios. You will also classify vulnerabilities, identify systemic risks, and create reproducible reports and attack cases that support model safety improvements.

  • Contractor position
  • Part-time engagement requiring 20 or more hours per week
  • Pay range: $48-$62 per hour
  • Expert-level opportunity
  • Text-based AI safety evaluation work

What You'll Do

You will conduct adversarial evaluations of conversational AI systems and document the results clearly for both technical and non-technical stakeholders.

  • Red team conversational AI systems with adversarial inputs and structured attack cases
  • Find jailbreak, prompt injection, and multi-turn manipulation failures
  • Annotate failures and classify vulnerabilities using established taxonomies
  • Flag systemic risks across model behaviors and evaluation scenarios
  • Produce reproducible vulnerability reports and attack cases
  • Test sensitive safety topics including bias, misinformation, and harmful behaviors
  • Work with frameworks, benchmarks, and playbooks to maintain consistent testing

Required Qualifications

This role is designed for an expert practitioner with demonstrated experience probing AI systems or related socio-technical and cybersecurity risks. Clear, precise communication is essential because findings must be understandable to varied audiences.

  • Prior AI red teaming or adversarial model testing experience
  • Experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to find jailbreak, prompt injection, and multi-turn manipulation failures
  • Experience writing reproducible vulnerability reports and attack cases
  • Strong ability to explain risks clearly to technical and non-technical stakeholders
  • Comfort using frameworks, taxonomies, benchmarks, and playbooks
  • Native fluency in English and Danish

Helpful Background

The following experience can support success in this role, although the required qualifications above define the core expectations.

  • Adversarial machine learning experience involving jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction
  • Cybersecurity experience in penetration testing, exploit development, or reverse engineering
  • Socio-technical risk experience involving harassment, disinformation, or abuse analysis
  • Creative probing skills developed through psychology, acting, or writing

Who Can Apply

This opportunity is available to applicants located in the eligible countries listed for the project: Austria, Belgium, Bulgaria, Canada, Cyprus, Czechia, Germany, Denmark, Estonia, Spain, Finland, France, the United Kingdom, Greece, Croatia, Hungary, Ireland, Italy, Lithuania, Luxembourg, Latvia, Malta, the Netherlands, Poland, Portugal, Romania, Sweden, Slovenia, Slovakia, and the United States.

The work is structured for a part-time contractor who can commit at least 20 hours each week and meets the English and Danish fluency requirement.

Build Your AI Training Career With OpenTrain

AI training and data labeling are part of the human side of building artificial intelligence. Contributors evaluate model outputs, prepare examples, and identify problems that help modern AI systems become more capable and reliable.

OpenTrain helps you build a durable portfolio of specialized AI work. Your profile can make it easier to show credible experience, find projects aligned with your expertise, and grow in this fast-moving field.

  • Create an OpenTrain account for free
  • Build a profile around your AI safety and evaluation experience
  • Discover specialized projects in one place
  • Apply in minutes and develop a long-term AI training portfolio

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

AI Safety Red Team Expert

Use structured red team methods to uncover vulnerabilities in conversational AI systems and agents. This part-time contractor role pays $29-$45 per hour and requires native English and Portuguese fluency.

Generative AI & RLHF
Text
Remote · Austria, Belgium, Bulgaria +27 more
English, Portuguese
Part-time · Flexible
Intermediate level
Hourly · $29–$45/hr

Posted Jul 30, 2026

AI Safety Red Team Expert

Probe conversational AI for jailbreaks, prompt injections, bias, and misuse as an expert red teamer. Work remotely as a Dutch and English native speaker for $48 to $62 per hour, with a 20+ hour weekly commitment.

Generative AI & RLHF
Text
Remote · Austria, Belgium, Bulgaria +27 more
English, Dutch
Part-time · Flexible
Expert level
Hourly · $48–$62/hr

Posted Jul 30, 2026

AI Safety Red Team Expert English Thai

Test conversational AI for jailbreaks, prompt injections, misuse, bias, and manipulation as an English and Thai AI Safety Red Team Expert. This remote contractor role offers $24-$35 per hour and about 40 hours per week.

Generative AI & RLHF
Text
Remote · Austria, Belgium, Bulgaria +27 more
Thai, English
Part-time · Flexible
Expert level
Hourly · $24–$35/hr

Posted Jul 30, 2026