Join OpenTrain AI as an expert red teamer probing conversational models for jailbreaks, prompt injection, bias exploitation, and multi-turn manipulation; remote contractor role, 20+ hrs/week, $48–$62/hr, native English and Dutch required.
Generative AI & RLHF
100% Remote Hourly · $48–$62/hr
$48–$62/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Jul 30, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover specialized projects, build a lasting portfolio of AI training work, and grow freelance careers that focus on the human side of AI.
OpenTrain AI is the hiring and contracting organization for this role — we post the work, provide the evaluation frameworks and playbooks you'll follow, and pay contractors directly.
About AI training and red teaming
AI training (often called data labeling or human feedback work) is the human-led work that teaches models how to behave. Red teaming is a specialized form of that work: adversarial testers probe models to surface vulnerabilities, misuse paths, and safety failures so engineers can fix them.
This role focuses on text-based red teaming for conversational AI: creating reproducible attack cases, annotating failures, and producing structured datasets and reports that drive safety improvements.
The role
You will be an AI Safety Red Teaming Expert conducting adversarial tests against conversational models and agents. Work is text-based and centers on surfacing jailbreaks, prompt injections, multi-turn manipulation, bias exploitation, and other misuse scenarios.
Tasks include classification and annotation (RLHF and evaluation-rating style work), documenting reproducible attack cases, and producing reports and datasets organized for engineers and product teams.
Data type: TEXT; label types: RLHF and EVALUATION_RATING
Employment: Contractor, Part-time
What you'll do
Red team conversational AI using adversarial prompts, roleplay, and multi-turn strategies to reveal vulnerabilities.
Identify failures, classify and document vulnerabilities, and flag systemic risks in outputs.
Follow taxonomies, benchmarks, and playbooks to keep evaluations consistent, reproducible, and actionable.
Produce structured reports, datasets, and attack cases that engineers and product teams can act on.
Requirements
This is an expert-level role. Preserve the required skills below exactly as you meet them — we will assess these in your application and sample tasks.
Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
Ability to test conversational AI for jailbreaks, prompt injection, misuse, and bias exploitation
Structured evaluation habits: use frameworks, taxonomies, or benchmarks for reproducible testing
Clear written communication explaining risks to technical and non-technical stakeholders
Native fluency in English and Dutch
Available for roughly 20+ hours per week
Helpful background
Adversarial ML experience with jailbreak datasets, prompt-injection or model-extraction attacks
Cybersecurity experience such as penetration testing, exploit development, or reverse engineering
Socio-technical risk experience like harassment, disinformation probing, or abuse analysis
Creative probing skills from psychology, acting, creative writing, or related fields
Compensation, schedule, and logistics
Pay is hourly on a contractor basis: $48–$62 per hour (OpenTrain AI pays contractors directly). The role is part-time and expects at least 20 hours per week; projects may allow flexible scheduling.
This role is open worldwide. You must be fluent in English and Dutch and able to work independently following structured playbooks and benchmarks.
Hourly rate: USD $48–$62 / hour
Time requirement: 20+ hours per week
Location: Remote, worldwide
How it works and how to apply
Create an OpenTrain account (free) to apply. Your profile showcases your experience, language skills, and past red teaming work so you can be matched to projects and tasks.
If selected you'll complete sample evaluations and follow provided playbooks and taxonomies. Deliverables typically include annotated examples, classified failure cases, and reproducible attack logs.
OpenTrain helps you build a persistent AI training portfolio that customers and projects can see
Applications will be evaluated for prior red teaming experience, structured testing ability, and language fluency
Join OpenTrain AI to red-team conversational models—finding jailbreaks, prompt injections, bias, and misuse—while producing reproducible attack cases and risk reports. Contract, part-time role (20+ hrs/week) paying USD 17–25/hr; native English and Indonesian required.
Join OpenTrain as an AI Safety Red Team Expert to probe conversational AI with adversarial prompts, jailbreaks, and structured red‑team methods. Remote, contract, 20+ hrs/week, pay $29–$45/hr; native English and Portuguese required.
Probe conversational AI for jailbreaks, prompt injections, and misuse as a remote AI Safety Red Team Specialist for OpenTrain. Contractor, part-time role (20+ hrs/week) paying $24–$35/hr; native English and Thai required.