Probe conversational AI models and agents for jailbreaks, prompt injections, bias, and other safety failures. This expert, remote contract role offers flexible part-time work at $48–$62 per hour for fluent English and Swedish speakers.
Generative AI & RLHF
100% Remote Hourly · $48–$62/hr
$48–$62/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Jul 31, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. OpenTrain helps people discover opportunities, build their AI training careers, and apply to projects in one place.
Remote contract work
Part-time schedule of 20+ hours per week
Pay of $48–$62 USD per hour
About AI Safety Red Teaming
AI training is the human side of building artificial intelligence. In safety red teaming, experts deliberately test models with adversarial prompts and realistic misuse scenarios to uncover weaknesses that automated evaluations may miss.
Your findings will help identify unsafe behaviors, clarify systemic risks, and improve how conversational AI models and agents respond in challenging situations.
Work directly with cutting-edge conversational AI systems
Use structured evaluation to surface and document model risks
Contribute to safer, more reliable AI behavior
The Role
OpenTrain AI is seeking an expert AI Safety Red Teamer for text-based adversarial testing of conversational AI models and agents. The work combines creative probing, structured evaluation, vulnerability classification, and clear reporting.
You will investigate sensitive areas including bias, misinformation, harmful behavior, and other misuse risks while following established guidance, taxonomies, benchmarks, and playbooks.
Experience level: Expert
Data type: Text
Languages: Native English and Swedish fluency
Employment type: Contractor, part-time
Worldwide opportunity
What You'll Do
You will design and execute adversarial tests, identify meaningful failure modes, and turn your observations into structured materials that can support model safety improvements.
Red team conversational AI models and agents with adversarial prompts and multi-turn tactics
Probe jailbreaks, prompt injections, misuse cases, bias exploitation, and manipulation
Annotate failures and classify vulnerabilities
Flag systemic risks in model behavior
Apply taxonomies, benchmarks, and playbooks to keep testing consistent
Document findings in reproducible reports, datasets, and attack cases
Work thoughtfully on sensitive topics such as bias, misinformation, and harmful behavior
Required Qualifications
This role requires prior hands-on experience with AI red teaming, adversarial model evaluation, adversarial AI, cybersecurity, or socio-technical probing. You should be able to think like an attacker while communicating findings precisely to both technical and non-technical stakeholders.
Prior experience in AI red teaming or adversarial model evaluation
Familiarity with jailbreaks, prompt injections, misuse cases, or multi-turn manipulation
Ability to identify failure modes that automated tests may miss
Ability to classify vulnerabilities and explain risks clearly
Native fluency in English and Swedish
Strong written communication
Comfort working with structured frameworks, taxonomies, benchmarks, or playbooks
Helpful Background
The following experience can strengthen your application and help you contribute across a wider range of adversarial testing scenarios.
Adversarial machine learning experience involving jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction
Cybersecurity experience in penetration testing, exploit development, or reverse engineering
Experience probing harassment, disinformation, abuse, or conversational AI risks
Creative and unconventional problem-solving in adversarial settings
Why Join AI Training Work
AI training and data labeling are growing fields where people help shape how modern AI systems behave. Contributors use their judgment to evaluate outputs, uncover risks, and prepare the human feedback that supports better models.
This remote, flexible format can fit around other commitments while giving experienced specialists a direct role in the development of advanced AI systems.
Work remotely from anywhere with a computer and internet connection
Choose flexible part-time work around your schedule
Apply specialist expertise to state-of-the-art AI evaluation
Build experience in a rapidly growing technology field
Stress-test frontier AI systems through adversarial prompt design, jailbreak testing, and safety evaluation. This expert contract role offers flexible, remote work at $70–$84 per hour for 20+ hours weekly across eligible countries.
Join OpenTrain as an AI Safety Red Teamer probing conversational models for jailbreaks, prompt injections, and misuse. Remote contractor role, 20+ hrs/week, $20–22/hr; native fluency in English and Odia required.
Probe conversational AI for jailbreaks, prompt injections, and misuse as a remote AI Safety Red Team Specialist for OpenTrain. Contractor, part-time role (20+ hrs/week) paying $24–$35/hr; native English and Thai required.