Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Team Specialist (Conversational Models)

Join OpenTrain as an AI Safety Red Team Specialist testing conversational models with adversarial prompts and jailbreaks; remote, 20+ hrs/week, $20–$22/hr. Requires expert red-teaming experience and native fluency in English and Punjabi.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $20–$22/hr

$20–$22/hr

Compensation

Worldwide

Eligibility

Expert

Experience

Jul 13, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for people building careers in AI training and data labeling. We connect independent contributors with specialized projects that directly shape how AI systems behave, while helping freelancers build a durable portfolio of evaluation work.

About AI training and red teaming

AI training (also called data labeling or human feedback work) is the human layer behind modern models — people prepare examples, evaluate outputs, and flag failures that automated tests miss. Red teaming is a specialized strand of that work: adversarially probing models to find jailbreaks, prompt-injection paths, bias, misinformation, and multi-turn manipulation.

This role focuses on text-based adversarial testing of conversational AI and producing reproducible attack cases, reports, and datasets that help safety teams fix vulnerabilities.

The role

You will work as an expert AI Safety Red Team Specialist running structured adversarial evaluations against conversational models and agents. The contract is part-time (20+ hours per week), remote, and open worldwide. You will document vulnerabilities, classify failure modes, and create reproducible materials for safety review.

  • Employment type: Contractor, Part-time
  • Time requirement: 20+ hours/week
  • Data type: Text; label types: RLHF, Evaluation Rating
  • Pay: $20–$22 USD per hour

What you'll do

Your day-to-day work is text-focused adversarial testing and clear documentation so engineering and safety teams can reproduce and remediate issues.

  • Design and execute adversarial prompts, jailbreaks, prompt-injection and multi-turn attack sequences following playbooks and taxonomies
  • Annotate failures and classify vulnerabilities using existing frameworks and benchmarks
  • Flag systemic risks and sensitive behavior (bias, misinformation, harmful instructions) while following safety guidelines
  • Produce reproducible reports, labeled datasets, and attack cases for model safety review and tracking

Requirements

We require demonstrable, expert-level experience and language skills. Preserve and reflect the structured testing habits you bring to every evaluation.

  • Prior red teaming or adversarial model-evaluation experience (AI adversarial work, cybersecurity, or socio-technical probing)
  • Familiarity with jailbreaks, prompt injections, misuse cases, and bias exploitation
  • Structured use of frameworks, taxonomies, benchmarks, or playbooks to ensure consistent testing
  • Clear written communication for technical and non-technical stakeholders
  • Native fluency in English (en) and Punjabi (pa)

Helpful background (not required but valued)

The following backgrounds will make you especially effective at this role and help you design creative, high-impact adversarial tests.

  • Adversarial ML experience (jailbreak datasets, prompt-injection research, RLHF/DPO attacks, model extraction)
  • Cybersecurity experience (penetration testing, exploit development, reverse engineering)
  • Experience probing harassment, disinformation, abuse, or other conversational AI risks
  • Creative adversarial thinking from psychology, acting, writing, or social engineering

How the project works, compensation, and next steps

OpenTrain hires contractors for part-time schedules and pays on an hourly basis. Work is performed remotely and may touch sensitive topics; you will follow provided safety and content guidelines when annotating and reporting. If you meet the requirements, apply with examples of prior red-team or adversarial evaluation work and a short note about how you structure tests and document reproducibility.

  • Compensation: $20–$22 USD per hour, paid per contractor agreement
  • Worldwide applicants accepted; you must be native-fluent in English and Punjabi
  • Labeling context: text-based RLHF and evaluation rating tasks; expect detailed reporting and dataset creation

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Safety Red Teaming Specialist

Join OpenTrain AI to adversarially test conversational models, surface vulnerabilities, and produce reproducible attack cases; remote contractor work paying $20–$22/hr for experienced red teamers fluent in English and Bengali.

Generative AI & RLHF
Text
Remote · Worldwide
English, Bangla
Part-time · Flexible
Expert level
Hourly · $20–$22/hr

Posted Jul 13, 2026

AI Safety Red Team Specialist, Remote

Join OpenTrain as an AI Safety Red Team Specialist to adversarially test conversational models, document reproducible jailbreaks and misuse cases, and help make AI safer. Part-time contract, 20+ hrs/week, $20–$22/hr; native fluency in English and Assamese required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Assamese
Part-time · Flexible
Entry level
Hourly · $20–$22/hr

Posted Jul 13, 2026

AI Safety Red Teamer (Conversational Model Tester)

Join OpenTrain as an AI Safety Red Teamer probing conversational models for jailbreaks, prompt injections, and misuse. Remote contractor role, 20+ hrs/week, $20–22/hr; native fluency in English and Odia required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Odia
Part-time · Flexible
Entry level
Hourly · $20–$22/hr

Posted Jul 13, 2026