Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Teamer

Probe conversational AI models and agents for jailbreaks, prompt injections, bias, and other safety failures. This expert, remote contract role offers flexible part-time work at $48–$62 per hour for fluent English and Swedish speakers.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $48–$62/hr

$48–$62/hr

Compensation

Worldwide

Eligibility

Expert

Experience

Jul 31, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. OpenTrain helps people discover opportunities, build their AI training careers, and apply to projects in one place.

  • Remote contract work
  • Part-time schedule of 20+ hours per week
  • Pay of $48–$62 USD per hour

About AI Safety Red Teaming

AI training is the human side of building artificial intelligence. In safety red teaming, experts deliberately test models with adversarial prompts and realistic misuse scenarios to uncover weaknesses that automated evaluations may miss.

Your findings will help identify unsafe behaviors, clarify systemic risks, and improve how conversational AI models and agents respond in challenging situations.

  • Work directly with cutting-edge conversational AI systems
  • Use structured evaluation to surface and document model risks
  • Contribute to safer, more reliable AI behavior

The Role

OpenTrain AI is seeking an expert AI Safety Red Teamer for text-based adversarial testing of conversational AI models and agents. The work combines creative probing, structured evaluation, vulnerability classification, and clear reporting.

You will investigate sensitive areas including bias, misinformation, harmful behavior, and other misuse risks while following established guidance, taxonomies, benchmarks, and playbooks.

  • Experience level: Expert
  • Data type: Text
  • Languages: Native English and Swedish fluency
  • Employment type: Contractor, part-time
  • Worldwide opportunity

What You'll Do

You will design and execute adversarial tests, identify meaningful failure modes, and turn your observations into structured materials that can support model safety improvements.

  • Red team conversational AI models and agents with adversarial prompts and multi-turn tactics
  • Probe jailbreaks, prompt injections, misuse cases, bias exploitation, and manipulation
  • Annotate failures and classify vulnerabilities
  • Flag systemic risks in model behavior
  • Apply taxonomies, benchmarks, and playbooks to keep testing consistent
  • Document findings in reproducible reports, datasets, and attack cases
  • Work thoughtfully on sensitive topics such as bias, misinformation, and harmful behavior

Required Qualifications

This role requires prior hands-on experience with AI red teaming, adversarial model evaluation, adversarial AI, cybersecurity, or socio-technical probing. You should be able to think like an attacker while communicating findings precisely to both technical and non-technical stakeholders.

  • Prior experience in AI red teaming or adversarial model evaluation
  • Familiarity with jailbreaks, prompt injections, misuse cases, or multi-turn manipulation
  • Ability to identify failure modes that automated tests may miss
  • Ability to classify vulnerabilities and explain risks clearly
  • Native fluency in English and Swedish
  • Strong written communication
  • Comfort working with structured frameworks, taxonomies, benchmarks, or playbooks

Helpful Background

The following experience can strengthen your application and help you contribute across a wider range of adversarial testing scenarios.

  • Adversarial machine learning experience involving jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction
  • Cybersecurity experience in penetration testing, exploit development, or reverse engineering
  • Experience probing harassment, disinformation, abuse, or conversational AI risks
  • Creative and unconventional problem-solving in adversarial settings

Why Join AI Training Work

AI training and data labeling are growing fields where people help shape how modern AI systems behave. Contributors use their judgment to evaluate outputs, uncover risks, and prepare the human feedback that supports better models.

This remote, flexible format can fit around other commitments while giving experienced specialists a direct role in the development of advanced AI systems.

  • Work remotely from anywhere with a computer and internet connection
  • Choose flexible part-time work around your schedule
  • Apply specialist expertise to state-of-the-art AI evaluation
  • Build experience in a rapidly growing technology field

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Safety Red Teamer

Stress-test frontier AI systems through adversarial prompt design, jailbreak testing, and safety evaluation. This expert contract role offers flexible, remote work at $70–$84 per hour for 20+ hours weekly across eligible countries.

Generative AI & RLHF
Text
Remote · United States, Denmark, Estonia +30 more
English
Part-time · Flexible
Expert level
Hourly · $70–$84/hr

Posted Jul 17, 2026

AI Safety Red Teamer (Conversational Model Tester)

Join OpenTrain as an AI Safety Red Teamer probing conversational models for jailbreaks, prompt injections, and misuse. Remote contractor role, 20+ hrs/week, $20–22/hr; native fluency in English and Odia required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Odia
Part-time · Flexible
Entry level
Hourly · $20–$22/hr

Posted Jul 13, 2026

AI Safety Red Team Specialist

Probe conversational AI for jailbreaks, prompt injections, and misuse as a remote AI Safety Red Team Specialist for OpenTrain. Contractor, part-time role (20+ hrs/week) paying $24–$35/hr; native English and Thai required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Thai
Part-time · Flexible
Intermediate level
Hourly · $24–$35/hr

Posted Jul 30, 2026