Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Teaming Specialist

Join OpenTrain as an expert AI Safety Red Teaming Specialist to probe conversational models for jailbreaks, prompt injections, and misuse; contractor, 20+ hrs/week, $48–$62/hr, remote, requires native English and Norwegian fluency.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $48–$62/hr

$48–$62/hr

Compensation

Worldwide

Eligibility

Expert

Experience

Jul 30, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the centralized platform where people build careers in AI training and data labeling. We connect expert contributors with high-impact projects, help you consolidate work across contracts, and let you grow a durable freelance career focused on making AI better and safer.

We hire and contract directly for this project — contributors are part of OpenTrain AI and work under our structured playbooks and evaluation taxonomies.

  • Work remotely and build an AI training portfolio you control.
  • Flexible, freelance-friendly schedules suitable for part-time or contractor work.
  • Projects focus on concrete, reproducible datasets and reports that improve model safety.

About AI training and red teaming

AI training (data labeling, annotation, and human feedback) is the human side of modern AI development. People create, test, and evaluate examples that teach models how to behave — and red teaming is the adversarial practice of finding where they fail.

As a red teamer you directly shape model behavior by surfacing vulnerabilities, producing attack cases, and helping engineering teams prioritize mitigations for real-world risk.

  • This is cutting-edge, impactful work that blends technical and socio-technical skills.
  • Many projects accept flexible hours and require only a computer and internet connection.

The role

OpenTrain is hiring AI Safety Red Teaming Specialists to probe conversational AI models and agents for jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.

This is a contractor, part-time role requiring 20+ hours per week. Pay is hourly at USD 48–62/hr (typical engagements pay within that band). Work is text-based and remote; selected contributors may engage with sensitive topics under clear guidelines.

  • Employment type: Contractor, Part-time.
  • Time commitment: 20+ hours/week.
  • Pay: USD 48–62 per hour, paid per hour.

What you'll do

Your day-to-day work focuses on adversarial exploration, structured evaluation, and clear reporting. You will follow playbooks and taxonomies to keep findings reproducible and actionable.

  • Red team conversational models with adversarial prompts and multi-turn attack strategies.
  • Identify and classify jailbreaks, prompt injections, misuse cases, bias, and other model weaknesses.
  • Annotate failures and flag systemic risks using provided taxonomies and benchmarks.
  • Produce reproducible attack cases, datasets, and written reports for engineering and policy teams.

Requirements

Candidates must meet the core skills and language requirements below. These are non-negotiable because the role requires precise adversarial thinking and bilingual reporting.

  • Prior experience in AI red teaming, adversarial ML, cybersecurity, or socio-technical probing.
  • Ability to identify jailbreaks, prompt injection attacks, model extraction attempts, and misuse cases.
  • Experience using taxonomies, benchmarks, or structured evaluation playbooks for reproducible results.
  • Strong structured thinking and clear written communication for both technical and non-technical audiences.
  • Native fluency in English and Norwegian (both required).

Helpful background

The following backgrounds are not required but often make contributors more effective and faster to ramp up on complex red teaming tasks.

  • Experience with jailbreak datasets, RLHF/DPO attack strategies, or prompt-injection research.
  • Penetration testing, exploit development, reverse engineering, or cybersecurity experience.
  • Prior work evaluating harassment, disinformation, or other socio-technical risk scenarios.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Safety Red Team Specialist

Probe conversational AI for jailbreaks, prompt injections, and misuse as a remote AI Safety Red Team Specialist for OpenTrain. Contractor, part-time role (20+ hrs/week) paying $24–$35/hr; native English and Thai required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Thai
Part-time · Flexible
Intermediate level
Hourly · $24–$35/hr

Posted Jul 30, 2026

AI Safety Red Teaming Specialist

Join OpenTrain AI to adversarially test conversational models, surface vulnerabilities, and produce reproducible attack cases; remote contractor work paying $20–$22/hr for experienced red teamers fluent in English and Bengali.

Generative AI & RLHF
Text
Remote · Worldwide
English, Bangla
Part-time · Flexible
Expert level
Hourly · $20–$22/hr

Posted Jul 13, 2026

AI Safety Red Team Specialist (Conversational Models)

Join OpenTrain as an AI Safety Red Team Specialist testing conversational models with adversarial prompts and jailbreaks; remote, 20+ hrs/week, $20–$22/hr. Requires expert red-teaming experience and native fluency in English and Punjabi.

Generative AI & RLHF
Text
Remote · Worldwide
English, Punjabi
Part-time · Flexible
Expert level
Hourly · $20–$22/hr

Posted Jul 13, 2026