Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Team Expert (EN + ID)

Join OpenTrain AI to red-team conversational models—finding jailbreaks, prompt injections, bias, and misuse—while producing reproducible attack cases and risk reports. Contract, part-time role (20+ hrs/week) paying USD 17–25/hr; native English and Indonesian required.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $17–$25/hr

$17–$25/hr

Compensation

Worldwide

Eligibility

Expert

Experience

Jul 30, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the centralized platform for people building careers in AI training and data labeling. We connect skilled contractors with structured, high-impact projects so contributors can track work, build a portfolio, and grow into lasting freelance careers.

OpenTrain AI is the hiring and contracting organization for this role. You will join a distributed community shaping how modern AI systems behave by doing the hands-on testing, annotation, and documentation that model developers rely on.

Why this AI training work matters

AI training (data labeling, annotation, and human feedback) is the human foundation of modern AI. Red teaming and adversarial evaluation directly improve safety and reliability by revealing real-world failure modes automatic tests miss.

This work is remote-friendly and often flexible, making it a strong fit for professionals who want part-time, impactful work that shapes state-of-the-art systems.

The role — AI Safety Red Team Expert

You will probe conversational models and agents with structured adversarial techniques: jailbreaks, prompt injections, multi-turn manipulation, bias exploitation, and other misuse cases. The work is text-based and requires accurately annotating failures, classifying vulnerabilities, and producing reproducible attack cases and reports.

You will follow taxonomies, benchmarks, and playbooks to keep testing consistent, and communicate findings clearly to both technical and non-technical stakeholders. Participation in higher-sensitivity review areas is optional and governed by explicit guidance.

What you'll do day-to-day

  • Red team conversational AI models and agents using adversarial inputs and structured attack methods.
  • Annotate model failures, label vulnerability types, and flag systemic risks following provided taxonomies.
  • Use established benchmarks and playbooks to keep tests reproducible and consistent across projects.
  • Produce reproducible reports, datasets, and attack cases that stakeholders can act on.
  • Work across moving projects and communicate findings clearly to technical and non-technical audiences.

Requirements

  • Prior red teaming or adversarial model-evaluation experience in AI safety, cybersecurity, or socio-technical probing.
  • Hands-on experience with jailbreaks, prompt injection, or misuse-case testing.
  • Comfort using frameworks, taxonomies, benchmarks, or playbooks rather than ad hoc testing.
  • Strong adversarial thinking: ability to uncover vulnerabilities automated tests miss.
  • Clear written communication for risk documentation and stakeholder handoff.
  • Native fluency in English and Indonesian (both required).

Helpful background (nice-to-have)

  • Adversarial machine learning experience (jailbreak datasets, prompt-injection, RLHF/DPO attack knowledge, or model extraction).
  • Cybersecurity skills such as penetration testing, exploit development, or reverse engineering.
  • Socio-technical risk analysis experience (harassment, disinformation, abuse probing).
  • Creative problem solving for unconventional adversarial thinking and multi-turn manipulation.

Schedule, pay, and logistics

This is a contract, part-time role with a time expectation of 20+ hours per week. Work is 100% remote and open worldwide.

Pay is hourly (PAY_PER_HOUR) in USD. The role lists an hourly range of USD 17–25/hr with a stated hourly rate up to USD 25/hr. OpenTrain handles contracting and payments.

  • Employment type: Contractor, Part-time.
  • Work type: Text-based annotation and evaluation (label types include RLHF, evaluation rating, and data collection).
  • Location: Remote, worldwide; native English and Indonesian required.

How the work is managed

You will receive structured playbooks, taxonomies, and benchmarks to guide testing and annotation. Tasks include producing labeled examples, filling structured forms, and writing short reproducible reports or attack case descriptions.

Expect iteration across moving projects: prioritization, feedback cycles, and regular communication with OpenTrain coordinators and reviewers. Optional higher-sensitivity tasks will be governed by explicit instructions and safeguards.

Who should apply

Apply if you are an experienced red teamer or adversarial evaluator who writes clearly, thinks like an attacker, and wants to shape model safety through concrete, reproducible test cases.

This role is ideal for people seeking flexible, part-time contract work that directly affects how conversational AI systems behave—especially bilingual English/Indonesian professionals with safety or cybersecurity backgrounds.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Safety Red Teaming Expert (English & Dutch)

Join OpenTrain AI as an expert red teamer probing conversational models for jailbreaks, prompt injection, bias exploitation, and multi-turn manipulation; remote contractor role, 20+ hrs/week, $48–$62/hr, native English and Dutch required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Dutch
Part-time · Flexible
Expert level
Hourly · $48–$62/hr

Posted Jul 30, 2026

AI Safety Red Teamer

Stress-test frontier AI systems through adversarial prompt design, jailbreak testing, and safety evaluation. This expert contract role offers flexible, remote work at $70–$84 per hour for 20+ hours weekly across eligible countries.

Generative AI & RLHF
Text
Remote · United States, Denmark, Estonia +30 more
English
Part-time · Flexible
Expert level
Hourly · $70–$84/hr

Posted Jul 17, 2026

AI Safety Red Team Specialist

Probe conversational AI for jailbreaks, prompt injections, and misuse as a remote AI Safety Red Team Specialist for OpenTrain. Contractor, part-time role (20+ hrs/week) paying $24–$35/hr; native English and Thai required.

Generative AI & RLHF
Text
Remote · Worldwide
English, Thai
Part-time · Flexible
Intermediate level
Hourly · $24–$35/hr

Posted Jul 30, 2026