Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Team Expert

Join OpenTrain to red team conversational AI—find jailbreaks, prompt injections, and multi-turn manipulation, and produce reproducible vulnerability reports; remote, contract, 20+ hrs/week, $48–$62/hr. Native English and Danish required.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $48–$62/hr

$48–$62/hr

Compensation

Worldwide

Eligibility

Expert

Experience

Jul 30, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the centralized platform where people start and grow careers in AI training and data labeling. We help freelancers discover specialized AI training projects, build a unified portfolio, and manage work that directly shapes how modern AI systems behave.

As the hiring organization for this role, OpenTrain coordinates projects, provides taxonomies and playbooks, and supports contributors turning AI training into durable freelance careers.

About AI training and red teaming work

AI training (data labeling, annotation, human-in-the-loop work) is the human side of building AI. Red teaming is a specialized subset: adversarially probing models to uncover jailbreaks, prompt injections, bias, misinformation, and other failure modes.

This role is text-focused, high-impact work: your findings feed benchmarks, playbooks, and model improvements used across teams building safer conversational AI.

The role

We’re hiring an AI Safety Red Team Expert to design and execute adversarial attacks against conversational models and agents, classify vulnerabilities, and produce reproducible attack cases and reports. You will follow structured taxonomies, benchmarks, and playbooks to keep testing consistent and actionable.

This is a remote, contract role for 20+ hours per week. Employment type: contractor, part-time. Work is paid hourly at $48–$62 USD per hour.

What you'll do

Focus on adversarial, multi-turn text interactions to reveal practical failure modes and systemic risks. Deliver clear, reproducible artifacts that engineers and safety teams can act on.

  • Red team conversational AI with adversarial inputs and structured attack cases.
  • Annotate failures and classify vulnerabilities using provided taxonomies and playbooks.
  • Flag systemic risks across bias, misinformation, harassment, and misuse scenarios.
  • Produce reproducible reports and attack cases that document steps, prompts, and model behavior.
  • Contribute to benchmarks and evaluation ratings tied to RLHF-style assessments.

Requirements

Candidates must meet the essential skills and availability below. OpenTrain preserves every project’s standards by requiring clear, reproducible reporting and adherence to playbooks.

  • Prior red teaming or adversarial model testing experience (AI, cybersecurity, or socio-technical probing).
  • Proven ability to find jailbreaks, prompt injection, and multi-turn manipulation failures.
  • Experience writing reproducible vulnerability reports and attack cases.
  • Comfortable using frameworks, taxonomies, benchmarks, and playbooks to keep testing consistent.
  • Strong ability to explain risks clearly to technical and non-technical stakeholders.
  • Native fluency in English and Danish.
  • Available for 20+ hours per week; contract, part-time engagement.

Helpful background

The following experience is not mandatory but will help you succeed and move faster in this role.

  • Adversarial ML or jailbreak dataset experience (prompt injection, model extraction, RLHF/DPO attacks).
  • Cybersecurity skills such as penetration testing, exploit development, or reverse engineering.
  • Socio-technical risk analysis experience (disinformation, harassment, abuse case analysis).
  • Creative probing skills from psychology, acting, or professional writing to craft realistic adversarial interactions.

How the work and compensation work

This project is text-based. Label types include RLHF-style evaluation ratings and other model-evaluation annotations. You will use structured forms and playbooks to record findings and produce reproducible examples.

Pay is hourly in USD: $48–$62/hr. You’ll be engaged as a contractor on a part-time schedule; OpenTrain supports remote contributors worldwide, but fluency requirements (English and Danish) apply.

  • Data type: text. Label types: RLHF, evaluation rating.
  • Employment: contractor, part-time. Time: 20+ hours/week.
  • Location: remote / worldwide, subject to language requirements.

Who should apply and next steps

Apply if you have hands-on red teaming or adversarial testing experience, enjoy creative probing, and can document reproducible attack cases clearly. This role is ideal for experienced practitioners who want to turn safety research and adversarial skills into reliable freelance work.

If you meet the language and experience requirements, create an OpenTrain profile, complete the required screening tasks, and submit examples of past red teaming or report samples when requested.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Safety Red Team Expert (Conversational Models)

Join OpenTrain as an AI Safety Red Team Expert to probe conversational AI with adversarial prompts, jailbreaks, and structured red‑team methods. Remote, contract, 20+ hrs/week, pay $29–$45/hr; native English and Portuguese required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Portuguese
Part-time · Flexible
Intermediate level
Hourly · $29–$45/hr

Posted Jul 30, 2026

AI Safety Red Team Expert (EN + ID)

Join OpenTrain AI to red-team conversational models—finding jailbreaks, prompt injections, bias, and misuse—while producing reproducible attack cases and risk reports. Contract, part-time role (20+ hrs/week) paying USD 17–25/hr; native English and Indonesian required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Indonesian
Part-time · Flexible
Expert level
Hourly · $17–$25/hr

Posted Jul 30, 2026

AI Safety Red Teaming Expert (English & Dutch)

Join OpenTrain AI as an expert red teamer probing conversational models for jailbreaks, prompt injection, bias exploitation, and multi-turn manipulation; remote contractor role, 20+ hrs/week, $48–$62/hr, native English and Dutch required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Dutch
Part-time · Flexible
Expert level
Hourly · $48–$62/hr

Posted Jul 30, 2026