Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Team Specialist

Probe conversational AI for jailbreaks, prompt injections, and misuse as a remote AI Safety Red Team Specialist for OpenTrain. Contractor, part-time role (20+ hrs/week) paying $24–$35/hr; native English and Thai required.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $24–$35/hr

$24–$35/hr

Compensation

Worldwide

Eligibility

Intermediate

Experience

Jul 30, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for people building careers in AI training and data labeling. We help freelancers find specialized projects, build a single verified portfolio of AI training work, and grow long-term freelance careers working on the human side of AI.

About AI training and red teaming

AI training (also called data labeling or human feedback) is how modern models learn from human examples and judgments. Red teaming is a specialized branch of this work: people craft adversarial prompts and multi-turn probes to expose vulnerabilities, bias, and misuse so models can be safer and more robust.

  • This role focuses on text-based adversarial testing, structured evaluation, and reproducible attack cases.
  • Work directly shapes how conversational AI handles sensitive topics and resists manipulation.

The role

You will work as an AI Safety Red Team Specialist conducting adversarial testing of conversational models and agents. This is remote contract work, part-time at 20+ hours per week, and pays between USD 24 and 35 per hour depending on task complexity and reviewer level.

  • Employment type: Contractor, Part-time.
  • Hours: 20+ hours/week (flexible scheduling).
  • Pay: $24–$35 per hour.
  • Work is text-based and focused on evaluation and RLHF-style tasks.

What you'll do

Use playbooks, taxonomies, and benchmarks to run repeatable red-team attacks; annotate and classify failures; and produce clear, reproducible reports and datasets that developers can use to fix issues.

  • Design adversarial prompts and multi-turn attack chains to find jailbreaks and prompt-injection vectors.
  • Annotate failures, classify vulnerability types, and flag systemic risks.
  • Follow structured taxonomies, benchmarks, and playbooks for consistent evaluation.
  • Produce reproducible attack cases, datasets, and concise reports.
  • Work on sensitive-topic review with provided guidance and safety support.

Requirements

We need testers who bring prior adversarial experience, strong judgment, and clear communication skills. The role expects intermediate-level experience and careful use of structured frameworks.

  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
  • Familiarity with jailbreaks, prompt injection, and misuse-case analysis.
  • Ability to classify vulnerabilities and explain risks to technical and non-technical stakeholders.
  • Comfort using taxonomies, benchmarks, and playbooks to keep testing consistent and reproducible.
  • Native-level fluency in English and Thai (required).

Who should apply

Apply if you have hands-on adversarial testing experience, enjoy investigative, structured evaluation work, and want flexible, remote contract work that directly improves model safety. This role suits people who can translate findings into clear, actionable reports.

  • Intermediate-level red teamers, cybersecurity professionals transitioning to AI safety, and experienced RLHF annotators.
  • People who prefer part-time, remote, project-based work and can commit 20+ hours/week.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Safety Red Teaming Specialist

Join OpenTrain as an expert AI Safety Red Teaming Specialist to probe conversational models for jailbreaks, prompt injections, and misuse; contractor, 20+ hrs/week, $48–$62/hr, remote, requires native English and Norwegian fluency.

Generative AI & RLHF
Text
Remote · Worldwide
English, Norwegian
Part-time · Flexible
Expert level
Hourly · $48–$62/hr

Posted Jul 30, 2026

AI Safety Red Teaming Specialist

Join OpenTrain AI to adversarially test conversational models, surface vulnerabilities, and produce reproducible attack cases; remote contractor work paying $20–$22/hr for experienced red teamers fluent in English and Bengali.

Generative AI & RLHF
Text
Remote · Worldwide
English, Bangla
Part-time · Flexible
Expert level
Hourly · $20–$22/hr

Posted Jul 13, 2026

AI Safety Red Team Specialist (Conversational Models)

Join OpenTrain as an AI Safety Red Team Specialist testing conversational models with adversarial prompts and jailbreaks; remote, 20+ hrs/week, $20–$22/hr. Requires expert red-teaming experience and native fluency in English and Punjabi.

Generative AI & RLHF
Text
Remote · Worldwide
English, Punjabi
Part-time · Flexible
Expert level
Hourly · $20–$22/hr

Posted Jul 13, 2026