AI Safety Red Team Expert
Test conversational AI for jailbreaks, prompt injections, bias, misuse, and multi-turn manipulation. This expert contract role offers 20+ hours weekly at $48-$62 per hour for native English and Dutch speakers.
Generative AI & RLHF
$48–$62/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Jul 30, 2026
Posted
Open worldwide
About OpenTrain
OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. Create a free profile to showcase your experience, discover projects that match your skills, and build a lasting portfolio in a fast-growing field.
- Remote AI training and data-labeling opportunities
- A profile for building and demonstrating your AI work
- Free account creation and streamlined applications
About AI Training Work
AI training is the human side of building artificial intelligence. Contributors test models, review outputs, create examples, and identify weaknesses so AI systems become more useful, reliable, and safe. This work puts experienced specialists close to the development of cutting-edge conversational AI.
- Help identify risks that automated tests may miss
- Work with text-based model evaluations and human feedback
- Apply structured judgment to emerging AI behaviors
The Role
OpenTrain AI is seeking an AI Safety Red Team Expert to probe conversational AI models and agents with adversarial inputs. You will investigate jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation while following clearly communicated content guidelines.
The work focuses on sensitive topics including misinformation, harmful behavior, and bias. You will turn overlooked weaknesses into reproducible evaluation artifacts that help technical teams understand and reduce systemic risk.
- Expert-level contract role
- Part-time engagement requiring 20+ hours per week
- Text-based conversational AI safety work
- Worldwide opportunity
- Pay of $48-$62 USD per hour
What You'll Do
You will design and execute adversarial evaluations across different models, scenarios, and risk areas. Your work will expand evaluation coverage beyond the failure modes automated testing can detect and produce clear evidence that supports safer AI development.
- Test conversational models and agents with jailbreaks and prompt injections
- Create misuse scenarios, bias exploitation tests, and multi-turn manipulation cases
- Annotate model failures and classify vulnerabilities
- Identify systemic risks across models and risk areas
- Use taxonomies, benchmarks, and playbooks to keep evaluations structured
- Create reproducible reports, datasets, and attack cases
- Communicate technical and non-technical risks clearly
Required Qualifications
You should bring prior experience with red teaming involving AI adversarial work, cybersecurity, or socio-technical probing. You must be able to push systems toward failure modes while maintaining a disciplined, structured testing approach.
Native fluency in both English and Dutch is required. You should also be adaptable when working across different models, scenarios, and risk areas.
- Prior AI red teaming, cybersecurity, or socio-technical probing experience
- Ability to design and execute jailbreak and prompt-injection tests
- Experience testing misuse, bias, and multi-turn manipulation scenarios
- Experience using taxonomies, benchmarks, or playbooks
- Ability to classify vulnerabilities and explain systemic risks
- Native fluency in English and Dutch
Helpful Background
Relevant experience may come from several technical, analytical, or creative disciplines. The strongest candidates can think adversarially, recognize subtle failure modes, and document findings in a way that supports consistent evaluation and practical risk reduction.
- Adversarial machine learning or jailbreak datasets
- Prompt injection, RLHF or DPO attacks, or model extraction
- Penetration testing, exploit development, or reverse engineering
- Abuse analysis, harassment probing, or misinformation probing
- Conversational AI testing
- Psychology, acting, or writing experience that supports creative adversarial thinking
Why This Work Matters
Your findings help strengthen AI systems by making overlooked vulnerabilities reproducible and easier to evaluate. Expanded coverage gives technical teams clearer insight into bias, misuse, harmful behavior, and other risks before they appear in production.
- Influence how conversational AI systems are evaluated for safety
- Turn discovered vulnerabilities into actionable evaluation artifacts
- Help reveal risks across automated and human testing gaps
How to Apply Through OpenTrain
Apply through OpenTrain AI to pursue this expert AI training contract. OpenTrain supports people building careers in data labeling and AI training by helping them present credible experience, find relevant work, and grow a portfolio over time.
- Create a free OpenTrain account
- Build a profile that reflects your red teaming and AI safety experience
- Apply in minutes and manage your AI training career in one place
Keep exploring
Explore related jobs
Browse related job pages
Expertise