Test conversational AI for jailbreaks, prompt injections, bias, misuse, and multi-turn manipulation. This expert contract role offers 20+ hours weekly at $48-$62 per hour for native English and Dutch speakers.
Generative AI & RLHF
100% Remote Hourly · $48–$62/hr
$48–$62/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Jul 30, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. Create a free profile to showcase your experience, discover projects that match your skills, and build a lasting portfolio in a fast-growing field.
Remote AI training and data-labeling opportunities
A profile for building and demonstrating your AI work
Free account creation and streamlined applications
About AI Training Work
AI training is the human side of building artificial intelligence. Contributors test models, review outputs, create examples, and identify weaknesses so AI systems become more useful, reliable, and safe. This work puts experienced specialists close to the development of cutting-edge conversational AI.
Help identify risks that automated tests may miss
Work with text-based model evaluations and human feedback
Apply structured judgment to emerging AI behaviors
The Role
OpenTrain AI is seeking an AI Safety Red Team Expert to probe conversational AI models and agents with adversarial inputs. You will investigate jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation while following clearly communicated content guidelines.
The work focuses on sensitive topics including misinformation, harmful behavior, and bias. You will turn overlooked weaknesses into reproducible evaluation artifacts that help technical teams understand and reduce systemic risk.
Expert-level contract role
Part-time engagement requiring 20+ hours per week
Text-based conversational AI safety work
Worldwide opportunity
Pay of $48-$62 USD per hour
What You'll Do
You will design and execute adversarial evaluations across different models, scenarios, and risk areas. Your work will expand evaluation coverage beyond the failure modes automated testing can detect and produce clear evidence that supports safer AI development.
Test conversational models and agents with jailbreaks and prompt injections
Create misuse scenarios, bias exploitation tests, and multi-turn manipulation cases
Annotate model failures and classify vulnerabilities
Identify systemic risks across models and risk areas
Use taxonomies, benchmarks, and playbooks to keep evaluations structured
Create reproducible reports, datasets, and attack cases
Communicate technical and non-technical risks clearly
Required Qualifications
You should bring prior experience with red teaming involving AI adversarial work, cybersecurity, or socio-technical probing. You must be able to push systems toward failure modes while maintaining a disciplined, structured testing approach.
Native fluency in both English and Dutch is required. You should also be adaptable when working across different models, scenarios, and risk areas.
Prior AI red teaming, cybersecurity, or socio-technical probing experience
Ability to design and execute jailbreak and prompt-injection tests
Experience testing misuse, bias, and multi-turn manipulation scenarios
Experience using taxonomies, benchmarks, or playbooks
Ability to classify vulnerabilities and explain systemic risks
Native fluency in English and Dutch
Helpful Background
Relevant experience may come from several technical, analytical, or creative disciplines. The strongest candidates can think adversarially, recognize subtle failure modes, and document findings in a way that supports consistent evaluation and practical risk reduction.
Adversarial machine learning or jailbreak datasets
Prompt injection, RLHF or DPO attacks, or model extraction
Penetration testing, exploit development, or reverse engineering
Abuse analysis, harassment probing, or misinformation probing
Conversational AI testing
Psychology, acting, or writing experience that supports creative adversarial thinking
Why This Work Matters
Your findings help strengthen AI systems by making overlooked vulnerabilities reproducible and easier to evaluate. Expanded coverage gives technical teams clearer insight into bias, misuse, harmful behavior, and other risks before they appear in production.
Influence how conversational AI systems are evaluated for safety
Turn discovered vulnerabilities into actionable evaluation artifacts
Help reveal risks across automated and human testing gaps
How to Apply Through OpenTrain
Apply through OpenTrain AI to pursue this expert AI training contract. OpenTrain supports people building careers in data labeling and AI training by helping them present credible experience, find relevant work, and grow a portfolio over time.
Create a free OpenTrain account
Build a profile that reflects your red teaming and AI safety experience
Apply in minutes and manage your AI training career in one place
Probe conversational AI for jailbreaks, prompt injections, bias exploitation, and manipulation as an expert red team contractor. Work worldwide for $48 to $62 per hour, 20+ hours weekly, using English and Danish.
Help improve conversational AI by uncovering jailbreaks, prompt injections, bias risks, and other vulnerabilities. This remote contractor role offers 20+ hours per week at $29-$45 per hour for native English and Portuguese speakers.
Probe conversational AI models and agents for vulnerabilities using adversarial testing, jailbreaks, prompt injections, and multi-turn manipulation. This remote, part-time expert contract pays $17 to $25 per hour.