Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Team Specialist, Text Adversarial Testing

OpenTrain is hiring an expert AI Safety Red Team Specialist to perform text-based adversarial testing of conversational models (20+ hrs/week). Earn $48–$62/hr while finding jailbreaks, prompt injections, and systemic failure cases; native English and Finnish required.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $48–$62/hr

$48–$62/hr

Compensation

Worldwide

Eligibility

Expert

Experience

Jul 30, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for people building careers in AI training and data labeling. We help contributors find specialist projects, consolidate experience across work, and build a reusable portfolio that demonstrates their expertise.

OpenTrain is hiring and contracting directly for this role. Creating an OpenTrain account is free and lets you apply, track work, and grow a durable freelance career in AI training.

Why AI Training and Red Team Work Matters

AI training is the human side of building intelligent systems — people prepare, test, and correct the examples models learn from. Red teaming is a high-impact niche: adversarial testing surfaces real-world failure modes that automated checks often miss.

This work is remote and flexible, often suitable as part-time or contract work. Contributors directly shape how state-of-the-art conversational AI behaves, improving safety, robustness, and user trust.

The Role

As an AI Safety Red Team Specialist you will perform text-based adversarial testing of conversational models and agents. Your output will include annotated failures, classified vulnerabilities, and reproducible attack cases that feed into model safety reviews and datasets.

This is a contractor, part-time role (20+ hours per week). Pay is hourly at $48–$62 USD per hour, paid per timesheet. Work is worldwide-remote; you must be natively fluent in English and Finnish.

  • Work type: Contractor, Part-time
  • Time commitment: 20+ hours/week
  • Pay: $48–$62 USD per hour
  • Languages required: Native English and Finnish
  • Data type: Text (RLHF and evaluation rating tasks)

What You'll Do

You will design, execute, and document adversarial test cases against conversational AI systems using multi-turn attacks and structured methodologies. Maintain reproducibility and consistency by applying taxonomies, benchmarks, and playbooks.

  • Red team conversational models with jailbreaks, prompt injections, misuse scenarios, and multi-turn manipulation.
  • Annotate failures and classify vulnerabilities using established taxonomies and benchmarks.
  • Flag systemic risks and prioritize issues by severity and exploitability.
  • Produce reproducible reports, labeled datasets, and attack cases for safety review and model improvement.

Requirements

You must be an experienced red teamer or safety evaluator with demonstrated ability to probe models and create high-quality, reproducible reporting. Preserve clear, structured evaluation habits and communicate findings to diverse stakeholders.

  • Prior red teaming experience in AI adversarial testing, cybersecurity, or socio-technical probing.
  • Demonstrated ability to test conversational models with jailbreaks, prompt injections, and misuse cases.
  • Strong written communication for clear technical and non-technical risk reporting.
  • Structured evaluation habits: use frameworks, taxonomies, or benchmarks rather than ad hoc testing.
  • Native fluency in English and Finnish.
  • Experience level: Expert

Who Should Apply

Apply if you are a security researcher, AI safety specialist, RLHF evaluator, or experienced red teamer who enjoys creative adversarial thinking and rigorous documentation. This role suits people who prefer remote, flexible contract work and want to help shape safer conversational AI.

  • Candidates with backgrounds in cybersecurity, red teaming, threat modeling, or RLHF evaluation are a strong fit.
  • Ideal for people who can turn adversarial tests into reproducible datasets and clear recommendations.

How It Works

Create a free OpenTrain account to apply. We'll evaluate your experience and may request examples of past red teaming or adversarial testing work. If selected, you will receive onboarding materials, taxonomies, and playbooks to ensure consistent evaluations.

Work is remote and tracked through OpenTrain. Payments follow the stated hourly range and are processed per contract terms. OpenTrain supports contributors as they grow their AI training careers.

  • Application: Free OpenTrain account and submission of experience/examples.
  • Onboarding: Playbooks, taxonomies, and evaluation guidelines provided.
  • Payments: Hourly at $48–$62 USD, contractor terms.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Safety Red Team Specialist (Conversational Models)

Join OpenTrain as an AI Safety Red Team Specialist testing conversational models with adversarial prompts and jailbreaks; remote, 20+ hrs/week, $20–$22/hr. Requires expert red-teaming experience and native fluency in English and Punjabi.

Generative AI & RLHF
Text
Remote · Worldwide
English, Punjabi
Part-time · Flexible
Expert level
Hourly · $20–$22/hr

Posted Jul 13, 2026

AI Safety Red Team Specialist

Probe conversational AI for jailbreaks, prompt injections, and misuse as a remote AI Safety Red Team Specialist for OpenTrain. Contractor, part-time role (20+ hrs/week) paying $24–$35/hr; native English and Thai required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Thai
Part-time · Flexible
Intermediate level
Hourly · $24–$35/hr

Posted Jul 30, 2026

AI Safety Red Teaming Specialist

Join OpenTrain AI to adversarially test conversational models, surface vulnerabilities, and produce reproducible attack cases; remote contractor work paying $20–$22/hr for experienced red teamers fluent in English and Bengali.

Generative AI & RLHF
Text
Remote · Worldwide
English, Bangla
Part-time · Flexible
Expert level
Hourly · $20–$22/hr

Posted Jul 13, 2026