Probe conversational AI models and agents for vulnerabilities using adversarial testing, jailbreaks, prompt injections, and multi-turn manipulation. This remote, part-time expert contract pays $17 to $25 per hour.
Generative AI & RLHF
100% Remote Hourly · $17–$25/hr
$17–$25/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Jul 30, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover specialized projects, build a portable AI training profile, and apply in minutes.
You will work remotely on flexible AI safety projects while contributing to a growing field that helps make advanced AI systems more robust, safe, and trustworthy.
Remote contractor opportunity
Part-time schedule of 20+ hours per week
Compensation of $17 to $25 USD per hour
Worldwide opportunity
About AI Safety Training
AI training is the human side of building artificial intelligence. People evaluate model behavior, prepare high-quality examples, identify failures, and provide structured feedback that helps AI systems perform more reliably.
Red teaming is a specialized form of AI evaluation. By probing models with adversarial inputs and documenting what goes wrong, contributors help identify vulnerabilities and risks before they affect real users.
Work directly with conversational AI models and agents
Create evaluation data that supports safer AI development
Use structured taxonomies, benchmarks, and playbooks
Contribute to cutting-edge AI safety work
The Role
OpenTrain is seeking an AI Safety Red Team Expert to conduct text-based safety evaluations of conversational AI models and agents. You will use an adversarial mindset to push systems toward their breaking points, generate high-quality red team data, and document vulnerabilities clearly.
The role requires native or bilingual fluency in English and Indonesian, prior red teaming experience, and the ability to communicate technical and non-technical risks in a clear, reproducible way.
Experience level: Expert
Languages: English and Indonesian
Data type: Text
Labeling focus: Evaluation, rating, red teaming, and classification
What You'll Do
You will investigate how conversational AI models and agents respond to adversarial scenarios, then turn your findings into consistent, actionable data and documentation.
Red team models and agents using jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
Annotate failures and classify vulnerabilities
Flag systemic risks revealed during testing
Follow taxonomies, benchmarks, and playbooks to keep evaluations consistent
Produce reproducible reports, datasets, and attack cases
Explain identified risks clearly to technical and non-technical stakeholders
Required Qualifications
This expert-level role is designed for candidates who have already worked in red teaming or a closely related area. You should be comfortable probing complex systems, applying structured evaluation methods, and adapting your approach across different projects and customers.
Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
Curiosity and an adversarial mindset focused on testing system limits
A structured approach using frameworks or benchmarks
Strong written and verbal communication skills
Adaptability across different projects and customers
Native or bilingual fluency in both English and Indonesian
Additional Experience That Helps
The following backgrounds are welcome and may help you approach adversarial testing from different angles. They are listed as nice-to-have experience rather than required qualifications.
Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, and model extraction
Cybersecurity experience in penetration testing, exploit development, or reverse engineering
Socio-technical risk work involving harassment or disinformation probing, abuse analysis, or conversational AI testing
Creative probing experience in psychology, acting, or writing for unconventional adversarial thinking
Why This Work Matters
Every major AI system depends on human evaluation and feedback. By uncovering failure modes and documenting them carefully, you will help shape how conversational AI behaves and support the development of safer, more trustworthy systems.
This work also offers a direct way to build experience in a fast-growing area of technology while contributing to projects that influence the future of AI.
Play a direct role in improving AI safety and robustness
Build specialized experience in AI red teaming
Work remotely with flexible scheduling
Contribute to a portable AI training career profile through OpenTrain
How to Apply
Create a free OpenTrain account and apply for this contractor opportunity in minutes. Be prepared to demonstrate your relevant red teaming experience, structured testing approach, communication skills, and English and Indonesian fluency.
Probe conversational AI with jailbreaks, prompt injections, and bias tests while documenting vulnerabilities and systemic risks. This remote expert contract requires native English and Malay fluency and offers $17 to $25 per hour.
Use expert adversarial testing to expose jailbreaks, prompt injections, bias, misinformation, and harmful behaviors in conversational AI. This remote contractor role offers $24-$35 per hour and about 40 hours per week.
Probe conversational AI for jailbreaks, prompt injections, bias exploitation, and manipulation as an expert red team contractor. Work worldwide for $48 to $62 per hour, 20+ hours weekly, using English and Danish.