Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Team Expert

Probe conversational AI for jailbreaks, prompt injections, bias exploitation, and harmful behavior. Use structured evaluations and reproducible reports to help strengthen AI safety, working 20+ hours per week at $16 to $22 per hour.

OpenTrain AI

Generative AI & RLHF

Remote Hourly · $16–$22/hr

$16–$22/hr

Compensation

30 countries

Eligibility

Entry

Experience

Jul 13, 2026

Posted

Open to applicants in

Austria Belgium Bulgaria
+27 more
  • Austria
  • Belgium
  • Bulgaria
  • Canada
  • Croatia
  • Cyprus
  • Czechia
  • Denmark
  • Estonia
  • Finland
  • France
  • Germany
  • Greece
  • Hungary
  • Ireland
  • Italy
  • Latvia
  • Lithuania
  • Luxembourg
  • Malta
  • Netherlands
  • Poland
  • Portugal
  • Romania
  • Slovakia
  • Slovenia
  • Spain
  • Sweden
  • United Kingdom
  • United States

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes.

Creating an OpenTrain account is free, and your profile can help you develop a lasting portfolio of experience in the rapidly growing AI training industry.

About AI Safety Training

AI training is the human side of building modern artificial intelligence. People review model behavior, create evaluation data, and identify failures that automated systems may not catch.

In this role, your safety-focused feedback will help expand how conversational AI systems are tested for harmful, biased, misleading, or manipulative behavior.

  • Remote contractor work with flexible scheduling
  • Part-time commitment of 20+ hours per week
  • Direct involvement in cutting-edge conversational AI evaluation

The Role

OpenTrain AI is seeking an AI Safety Red Team Expert to probe conversational AI models and agents with adversarial inputs. You will test how systems respond to jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.

You will review text-based outputs involving sensitive topics such as bias, misinformation, and harmful behavior. Your work will generate structured human data and reproducible artifacts that help strengthen AI systems and broaden their safety evaluation coverage.

  • Role type: Part-time contractor
  • Experience level: Entry level
  • Pay: $16 to $22 USD per hour
  • Data type: Text

What You'll Do

You will apply evaluation taxonomies, benchmarks, playbooks, and quality standards consistently across projects and task types. Your findings should be clear enough for both technical and non-technical audiences to understand and act on.

  • Probe conversational AI models and agents for jailbreaks and prompt injections
  • Test misuse cases, bias exploitation, and multi-turn manipulation
  • Annotate model failures and classify vulnerabilities
  • Identify systemic risks, subtle errors, inconsistencies, and gaps
  • Review whether responses are accurate, complete, appropriate, or harmful
  • Produce reproducible reports, datasets, and attack cases
  • Identify vulnerabilities that automated tests may miss
  • Help broaden scenario coverage across evaluation tasks

Requirements

This is an entry-level opportunity requiring strong judgment, careful reasoning, and the ability to follow structured guidance. Native fluency in both English and Assamese is required for nuanced language and content review.

  • Native fluency in English and Assamese
  • Strong judgment when assessing AI responses for accuracy, completeness, appropriateness, and harm
  • Ability to identify and classify vulnerabilities, subtle errors, inconsistencies, and systemic risks
  • Consistent application of taxonomies, benchmarks, playbooks, and quality standards
  • Clear communication with technical and non-technical audiences
  • Adaptability across projects, task types, and evaluation scenarios

Helpful Backgrounds

Relevant experience may include adversarial machine learning, jailbreak datasets, prompt injection, model extraction, cybersecurity, penetration testing, exploit development, reverse engineering, socio-technical risk, misinformation probing, abuse analysis, or conversational AI testing.

Backgrounds in psychology, acting, or writing may also support the unconventional adversarial thinking needed to explore how models can be manipulated.

  • Adversarial machine learning or conversational AI testing
  • Cybersecurity, penetration testing, exploit development, or reverse engineering
  • Prompt injection, jailbreak datasets, or model extraction
  • Socio-technical risk, misinformation probing, or abuse analysis
  • Psychology, acting, or writing

Eligibility and Work Details

This contractor opportunity is available to applicants located in the following countries: Austria, Belgium, Bulgaria, Canada, Cyprus, Czech Republic, Germany, Denmark, Estonia, Spain, Finland, France, United Kingdom, Greece, Croatia, Hungary, Ireland, Italy, Lithuania, Luxembourg, Latvia, Malta, Netherlands, Poland, Portugal, Romania, Sweden, Slovenia, Slovakia, and the United States.

  • Minimum time requirement: 20+ hours per week
  • Hourly compensation: $16 to $22 USD
  • Languages used: English and Assamese
  • Contract arrangement: Part time

Build Your AI Training Career

AI training and data-labeling work is one of the fastest-growing ways to participate in technology. Contributors help shape how state-of-the-art systems understand information, respond to people, and handle difficult real-world situations.

Through OpenTrain, you can build a profile around your AI training experience, discover projects aligned with your skills, and apply to opportunities in minutes.

  • Create a free OpenTrain account
  • Showcase your evaluation and safety experience
  • Apply for AI training opportunities through OpenTrain

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

AI Safety Red Team Expert

Use structured red team methods to uncover vulnerabilities in conversational AI systems and agents. This part-time contractor role pays $29-$45 per hour and requires native English and Portuguese fluency.

Generative AI & RLHF
Text
Remote · Austria, Belgium, Bulgaria +27 more
English, Portuguese
Part-time · Flexible
Intermediate level
Hourly · $29–$45/hr

Posted Jul 30, 2026

AI Safety Red Team Expert

Probe conversational AI for jailbreaks, prompt injections, bias, and misuse as an expert red teamer. Work remotely as a Dutch and English native speaker for $48 to $62 per hour, with a 20+ hour weekly commitment.

Generative AI & RLHF
Text
Remote · Austria, Belgium, Bulgaria +27 more
English, Dutch
Part-time · Flexible
Expert level
Hourly · $48–$62/hr

Posted Jul 30, 2026

AI Safety Red Team Expert

Test conversational AI for jailbreaks, prompt injections, bias exploitation, and multi-turn manipulation. This expert contract role offers $48-$62 per hour and requires native English and Danish fluency.

Generative AI & RLHF
Text
Remote · Austria, Belgium, Bulgaria +27 more
English, Danish
Part-time · Flexible
Expert level
Hourly · $48–$62/hr

Posted Jul 30, 2026