Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Teaming Expert (English & Dutch)

Join OpenTrain AI as an expert red teamer probing conversational models for jailbreaks, prompt injection, bias exploitation, and multi-turn manipulation; remote contractor role, 20+ hrs/week, $48–$62/hr, native English and Dutch required.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $48–$62/hr

$48–$62/hr

Compensation

Worldwide

Eligibility

Expert

Experience

Jul 30, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover specialized projects, build a lasting portfolio of AI training work, and grow freelance careers that focus on the human side of AI.

OpenTrain AI is the hiring and contracting organization for this role — we post the work, provide the evaluation frameworks and playbooks you'll follow, and pay contractors directly.

About AI training and red teaming

AI training (often called data labeling or human feedback work) is the human-led work that teaches models how to behave. Red teaming is a specialized form of that work: adversarial testers probe models to surface vulnerabilities, misuse paths, and safety failures so engineers can fix them.

This role focuses on text-based red teaming for conversational AI: creating reproducible attack cases, annotating failures, and producing structured datasets and reports that drive safety improvements.

The role

You will be an AI Safety Red Teaming Expert conducting adversarial tests against conversational models and agents. Work is text-based and centers on surfacing jailbreaks, prompt injections, multi-turn manipulation, bias exploitation, and other misuse scenarios.

Tasks include classification and annotation (RLHF and evaluation-rating style work), documenting reproducible attack cases, and producing reports and datasets organized for engineers and product teams.

  • Data type: TEXT; label types: RLHF and EVALUATION_RATING
  • Employment: Contractor, Part-time

What you'll do

  • Red team conversational AI using adversarial prompts, roleplay, and multi-turn strategies to reveal vulnerabilities.
  • Identify failures, classify and document vulnerabilities, and flag systemic risks in outputs.
  • Follow taxonomies, benchmarks, and playbooks to keep evaluations consistent, reproducible, and actionable.
  • Produce structured reports, datasets, and attack cases that engineers and product teams can act on.

Requirements

This is an expert-level role. Preserve the required skills below exactly as you meet them — we will assess these in your application and sample tasks.

  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to test conversational AI for jailbreaks, prompt injection, misuse, and bias exploitation
  • Structured evaluation habits: use frameworks, taxonomies, or benchmarks for reproducible testing
  • Clear written communication explaining risks to technical and non-technical stakeholders
  • Native fluency in English and Dutch
  • Available for roughly 20+ hours per week

Helpful background

  • Adversarial ML experience with jailbreak datasets, prompt-injection or model-extraction attacks
  • Cybersecurity experience such as penetration testing, exploit development, or reverse engineering
  • Socio-technical risk experience like harassment, disinformation probing, or abuse analysis
  • Creative probing skills from psychology, acting, creative writing, or related fields

Compensation, schedule, and logistics

Pay is hourly on a contractor basis: $48–$62 per hour (OpenTrain AI pays contractors directly). The role is part-time and expects at least 20 hours per week; projects may allow flexible scheduling.

This role is open worldwide. You must be fluent in English and Dutch and able to work independently following structured playbooks and benchmarks.

  • Hourly rate: USD $48–$62 / hour
  • Time requirement: 20+ hours per week
  • Location: Remote, worldwide

How it works and how to apply

Create an OpenTrain account (free) to apply. Your profile showcases your experience, language skills, and past red teaming work so you can be matched to projects and tasks.

If selected you'll complete sample evaluations and follow provided playbooks and taxonomies. Deliverables typically include annotated examples, classified failure cases, and reproducible attack logs.

  • OpenTrain helps you build a persistent AI training portfolio that customers and projects can see
  • Applications will be evaluated for prior red teaming experience, structured testing ability, and language fluency

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Safety Red Team Expert (EN + ID)

Join OpenTrain AI to red-team conversational models—finding jailbreaks, prompt injections, bias, and misuse—while producing reproducible attack cases and risk reports. Contract, part-time role (20+ hrs/week) paying USD 17–25/hr; native English and Indonesian required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Indonesian
Part-time · Flexible
Expert level
Hourly · $17–$25/hr

Posted Jul 30, 2026

AI Safety Red Team Expert (Conversational Models)

Join OpenTrain as an AI Safety Red Team Expert to probe conversational AI with adversarial prompts, jailbreaks, and structured red‑team methods. Remote, contract, 20+ hrs/week, pay $29–$45/hr; native English and Portuguese required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Portuguese
Part-time · Flexible
Intermediate level
Hourly · $29–$45/hr

Posted Jul 30, 2026

AI Safety Red Team Specialist

Probe conversational AI for jailbreaks, prompt injections, and misuse as a remote AI Safety Red Team Specialist for OpenTrain. Contractor, part-time role (20+ hrs/week) paying $24–$35/hr; native English and Thai required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Thai
Part-time · Flexible
Intermediate level
Hourly · $24–$35/hr

Posted Jul 30, 2026