Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Team Expert

Test conversational AI for jailbreaks, prompt injections, bias, misuse, and multi-turn manipulation. This expert contract role offers 20+ hours weekly at $48-$62 per hour for native English and Dutch speakers.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $48–$62/hr

$48–$62/hr

Compensation

Worldwide

Eligibility

Expert

Experience

Jul 30, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. Create a free profile to showcase your experience, discover projects that match your skills, and build a lasting portfolio in a fast-growing field.

  • Remote AI training and data-labeling opportunities
  • A profile for building and demonstrating your AI work
  • Free account creation and streamlined applications

About AI Training Work

AI training is the human side of building artificial intelligence. Contributors test models, review outputs, create examples, and identify weaknesses so AI systems become more useful, reliable, and safe. This work puts experienced specialists close to the development of cutting-edge conversational AI.

  • Help identify risks that automated tests may miss
  • Work with text-based model evaluations and human feedback
  • Apply structured judgment to emerging AI behaviors

The Role

OpenTrain AI is seeking an AI Safety Red Team Expert to probe conversational AI models and agents with adversarial inputs. You will investigate jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation while following clearly communicated content guidelines.

The work focuses on sensitive topics including misinformation, harmful behavior, and bias. You will turn overlooked weaknesses into reproducible evaluation artifacts that help technical teams understand and reduce systemic risk.

  • Expert-level contract role
  • Part-time engagement requiring 20+ hours per week
  • Text-based conversational AI safety work
  • Worldwide opportunity
  • Pay of $48-$62 USD per hour

What You'll Do

You will design and execute adversarial evaluations across different models, scenarios, and risk areas. Your work will expand evaluation coverage beyond the failure modes automated testing can detect and produce clear evidence that supports safer AI development.

  • Test conversational models and agents with jailbreaks and prompt injections
  • Create misuse scenarios, bias exploitation tests, and multi-turn manipulation cases
  • Annotate model failures and classify vulnerabilities
  • Identify systemic risks across models and risk areas
  • Use taxonomies, benchmarks, and playbooks to keep evaluations structured
  • Create reproducible reports, datasets, and attack cases
  • Communicate technical and non-technical risks clearly

Required Qualifications

You should bring prior experience with red teaming involving AI adversarial work, cybersecurity, or socio-technical probing. You must be able to push systems toward failure modes while maintaining a disciplined, structured testing approach.

Native fluency in both English and Dutch is required. You should also be adaptable when working across different models, scenarios, and risk areas.

  • Prior AI red teaming, cybersecurity, or socio-technical probing experience
  • Ability to design and execute jailbreak and prompt-injection tests
  • Experience testing misuse, bias, and multi-turn manipulation scenarios
  • Experience using taxonomies, benchmarks, or playbooks
  • Ability to classify vulnerabilities and explain systemic risks
  • Native fluency in English and Dutch

Helpful Background

Relevant experience may come from several technical, analytical, or creative disciplines. The strongest candidates can think adversarially, recognize subtle failure modes, and document findings in a way that supports consistent evaluation and practical risk reduction.

  • Adversarial machine learning or jailbreak datasets
  • Prompt injection, RLHF or DPO attacks, or model extraction
  • Penetration testing, exploit development, or reverse engineering
  • Abuse analysis, harassment probing, or misinformation probing
  • Conversational AI testing
  • Psychology, acting, or writing experience that supports creative adversarial thinking

Why This Work Matters

Your findings help strengthen AI systems by making overlooked vulnerabilities reproducible and easier to evaluate. Expanded coverage gives technical teams clearer insight into bias, misuse, harmful behavior, and other risks before they appear in production.

  • Influence how conversational AI systems are evaluated for safety
  • Turn discovered vulnerabilities into actionable evaluation artifacts
  • Help reveal risks across automated and human testing gaps

How to Apply Through OpenTrain

Apply through OpenTrain AI to pursue this expert AI training contract. OpenTrain supports people building careers in data labeling and AI training by helping them present credible experience, find relevant work, and grow a portfolio over time.

  • Create a free OpenTrain account
  • Build a profile that reflects your red teaming and AI safety experience
  • Apply in minutes and manage your AI training career in one place

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.