Use expert adversarial testing to expose jailbreaks, prompt injections, bias, misinformation, and harmful behaviors in conversational AI. This remote contractor role offers $24-$35 per hour and about 40 hours per week.
Generative AI & RLHF
100% Remote Hourly · $24–$35/hr
$24–$35/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Jul 30, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We connect skilled contributors with specialized projects, help them build a professional AI training profile, and make it easier to grow this work into a durable career. Creating an OpenTrain account is free.
Manage your AI training experience in one professional profile.
Discover projects that match your technical and language skills.
Apply to opportunities in minutes and build a credible portfolio over time.
About AI Safety Red Teaming
AI training is the human side of building artificial intelligence. Red team experts challenge conversational models with realistic and adversarial scenarios, helping identify weaknesses that automated tests may miss. Their findings become structured human data used to evaluate and strengthen AI systems.
This work sits at the cutting edge of AI safety, covering risks such as manipulation, misinformation, bias, abuse, and harmful model behavior. The role is text-based, remote, and focused on careful human judgment.
Work remotely with a computer and internet connection.
Help shape how conversational AI systems respond in difficult situations.
Participation in higher-sensitivity projects is optional.
The Role
OpenTrain is seeking an AI Safety Red Team Expert who is fluent in English and Thai. You will test conversational AI models and agents through adversarial inputs, investigate how systems fail, and convert your findings into structured evaluations and actionable reports.
This is remote, hourly contractor work at about 40 hours per week, with a minimum time expectation of 20 hours per week.
Experience level: Expert
Work arrangement: Remote, worldwide
Engagement: Hourly contractor and part-time
Rate: $24-$35 per hour
Primary data type: Text
Languages: English and Thai
What You'll Do
You will combine creative adversarial thinking with disciplined, repeatable testing. Your work will help expand evaluation coverage across misuse, manipulation, safety, and broader socio-technical risk scenarios.
Probe conversational AI models and agents with adversarial scenarios that automated testing may miss.
Annotate failures, classify vulnerabilities, and flag systemic risks in high-quality human data.
Apply taxonomies, benchmarks, frameworks, and playbooks to keep evaluations consistent and reproducible.
Document attack cases, datasets, and reports that communicate actionable findings.
Review outputs involving bias, misinformation, harmful behavior, and abuse with strong professional judgment.
Required Qualifications
This role requires prior experience in AI red teaming, cybersecurity, or socio-technical probing. You should be comfortable pushing systems toward failure points while maintaining consistent testing practices and explaining risks to both technical and non-technical audiences.
Prior AI red teaming or cybersecurity experience.
Native fluency in both English and Thai.
Ability to identify jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
Experience applying taxonomies, benchmarks, frameworks, or playbooks to adversarial testing.
Ability to communicate technical and socio-technical risks clearly.
Strong judgment when evaluating bias, misinformation, harmful behavior, and abuse.
Helpful Backgrounds
The following experience can support creative, technically informed probing. It is helpful but not listed as a required qualification.
Jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
Penetration testing, exploit development, or reverse engineering.
Harassment or disinformation probing, abuse analysis, or conversational AI testing.
Psychology, acting, or unconventional adversarial writing.
Sensitive Topic Support
Some evaluations may involve higher-sensitivity topics connected to harmful behavior, abuse, bias, or misinformation. Participation in these projects is optional, with clear topic guidance and wellness resources available.
Choose whether to participate in higher-sensitivity project work.
Follow project guidance when reviewing difficult or potentially harmful content.
Use structured evaluation practices to keep findings clear and professional.
How to Apply Through OpenTrain
Create a free OpenTrain account, build a profile that reflects your red teaming, cybersecurity, language, and evaluation experience, and apply in minutes. Your profile can help demonstrate relevant proof of work as you pursue specialized AI training opportunities.
Highlight English and Thai fluency.
Describe relevant adversarial testing, cybersecurity, and AI safety experience.
Show familiarity with frameworks, benchmarks, taxonomies, or testing playbooks.
Apply for this remote contractor opportunity through OpenTrain.
Probe conversational AI models and agents for vulnerabilities using adversarial testing, jailbreaks, prompt injections, and multi-turn manipulation. This remote, part-time expert contract pays $17 to $25 per hour.
Probe conversational AI with jailbreaks, prompt injections, and bias tests while documenting vulnerabilities and systemic risks. This remote expert contract requires native English and Malay fluency and offers $17 to $25 per hour.
Probe conversational AI for jailbreaks, prompt injections, bias exploitation, and manipulation as an expert red team contractor. Work worldwide for $48 to $62 per hour, 20+ hours weekly, using English and Danish.