Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Teaming Expert

Test conversational AI with jailbreaks, prompt injections, and other adversarial techniques while documenting vulnerabilities that improve model safety. This remote, part-time contractor role pays $16-$22 per hour and requires native English and Urdu fluency.

OpenTrain AI

Generative AI & RLHF

Remote Hourly · $16–$22/hr

$16–$22/hr

Compensation

30 countries

Eligibility

Entry

Experience

Jul 24, 2026

Posted

Open to applicants in

Austria Belgium Bulgaria
+27 more
  • Austria
  • Belgium
  • Bulgaria
  • Canada
  • Croatia
  • Cyprus
  • Czechia
  • Denmark
  • Estonia
  • Finland
  • France
  • Germany
  • Greece
  • Hungary
  • Ireland
  • Italy
  • Latvia
  • Lithuania
  • Luxembourg
  • Malta
  • Netherlands
  • Poland
  • Portugal
  • Romania
  • Slovakia
  • Slovenia
  • Spain
  • Sweden
  • United Kingdom
  • United States

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. Create a free profile to discover cutting-edge projects, showcase your experience, and apply in minutes.

  • Build a lasting portfolio of AI training and safety work
  • Find projects that match your skills and language experience
  • Work remotely as an independent contractor

About AI Safety Training

AI training is the human side of building artificial intelligence. Contributors test systems, evaluate model outputs, and prepare structured data that helps AI become more accurate, useful, and safe.

Red teaming is a specialized form of model evaluation. Human testers deliberately probe conversational AI with adversarial scenarios to uncover weaknesses that automated checks may miss.

  • Help shape how conversational AI handles challenging situations
  • Contribute to safer systems through structured human feedback
  • Work remotely with flexible scheduling around a 20-plus-hour weekly commitment

The Role

OpenTrain is seeking an AI Safety Red Teaming Expert to test conversational AI models and agents with adversarial inputs. You will identify weaknesses, classify vulnerabilities, and create structured human data that supports safer AI systems.

This text-based contractor role is listed as entry level and is available part time at $16 to $22 per hour. Higher-sensitivity assignments are optional and include clear topic guidance and wellness resources.

  • Role type: Part-time contractor
  • Experience level: Entry level
  • Time commitment: 20 or more hours per week
  • Pay: $16-$22 per hour
  • Data type: Text
  • Work format: Remote

What You'll Do

You will conduct structured adversarial testing and turn your findings into clear, reproducible documentation. The work includes evaluating model behavior across safety, accuracy, appropriateness, and completeness criteria.

  • Test conversational AI models and agents with jailbreaks and prompt injections
  • Probe misuse cases, bias exploitation, sensitive topics, and multi-turn manipulation
  • Annotate model failures and classify vulnerabilities
  • Flag systemic risks and identify gaps in evaluation coverage
  • Apply taxonomies, benchmarks, playbooks, and quality standards consistently
  • Document reproducible reports, datasets, and attack cases
  • Identify vulnerabilities that automated testing may miss
  • Collect and structure human data for AI safety improvements

Requirements

Native fluency in both English and Urdu is required. You should be able to make careful judgments about AI responses and explain your reasoning clearly to technical and non-technical audiences.

The role requires consistent, structured work against defined guidelines and quality standards. You should be comfortable noticing subtle errors, inconsistencies, unsafe behavior, and incomplete or inappropriate responses.

  • Native fluency in English and Urdu
  • Strong judgment when evaluating AI responses for accuracy, completeness, appropriateness, and safety
  • Ability to identify jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
  • Consistent use of taxonomies, benchmarks, playbooks, and quality standards
  • Ability to explain vulnerability assessments and reproducible attack cases clearly
  • Careful attention to subtle errors and inconsistencies

Helpful Backgrounds

Relevant experience can help you approach model testing from different angles, although these backgrounds are described as helpful rather than required. Unconventional adversarial thinking and strong written reasoning are valuable in this work.

  • Adversarial machine learning, including jailbreak datasets or prompt injection
  • RLHF or DPO attacks
  • Model extraction
  • Cybersecurity, penetration testing, exploit development, or reverse engineering
  • Socio-technical risk or abuse analysis
  • Harassment or misinformation probing
  • Conversational AI testing
  • Psychology, acting, or writing

Location And Language Eligibility

This project requires English and Urdu fluency and is available to contractors in the following countries.

  • Austria, Belgium, Bulgaria, Canada, Cyprus, Czechia, Germany, Denmark, Estonia, Spain, Finland, France, United Kingdom, Greece, Croatia, Hungary, Ireland, Italy, Lithuania, Luxembourg, Latvia, Malta, Netherlands, Poland, Portugal, Romania, Sweden, Slovenia, Slovakia, and the United States

Why Build Your AI Training Career With OpenTrain

AI safety work sits at the forefront of a rapidly growing technology industry. Through OpenTrain, you can develop a profile that shows credible experience, find projects aligned with your abilities, and build a career portfolio in AI training and data labeling.

  • Create an OpenTrain account for free
  • Apply to projects in minutes
  • Showcase your AI safety and evaluation experience
  • Grow from individual projects toward a durable AI training portfolio

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

AI Safety Red Team Expert

Use structured red team methods to uncover vulnerabilities in conversational AI systems and agents. This part-time contractor role pays $29-$45 per hour and requires native English and Portuguese fluency.

Generative AI & RLHF
Text
Remote · Austria, Belgium, Bulgaria +27 more
English, Portuguese
Part-time · Flexible
Intermediate level
Hourly · $29–$45/hr

Posted Jul 30, 2026

AI Safety Red Team Expert

Probe conversational AI for jailbreaks, prompt injections, bias, and misuse as an expert red teamer. Work remotely as a Dutch and English native speaker for $48 to $62 per hour, with a 20+ hour weekly commitment.

Generative AI & RLHF
Text
Remote · Austria, Belgium, Bulgaria +27 more
English, Dutch
Part-time · Flexible
Expert level
Hourly · $48–$62/hr

Posted Jul 30, 2026

AI Safety Red Team Expert

Test conversational AI for jailbreaks, prompt injections, bias exploitation, and multi-turn manipulation. This expert contract role offers $48-$62 per hour and requires native English and Danish fluency.

Generative AI & RLHF
Text
Remote · Austria, Belgium, Bulgaria +27 more
English, Danish
Part-time · Flexible
Expert level
Hourly · $48–$62/hr

Posted Jul 30, 2026