Join OpenTrain to perform adversarial, text-based red teaming on conversational AI—find jailbreaks, annotate failures, and produce reproducible attack cases. Remote contractor role paying $17–$25/hr that requires native English and Vietnamese and a background in red teaming or cybersecurity.
Generative AI & RLHF
100% Remote Hourly · $17–$25/hr
$17–$25/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Jul 30, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people starting and growing careers in AI training and data labeling. We help freelancers discover specialized AI training opportunities, build a unified portfolio of proof-of-work, and manage work across projects from a single place.
Why AI training matters
AI training (data labeling and human feedback) is the human side of building modern AI systems: people create examples, review outputs, and teach models how to behave. This work is remote, flexible, and accessible—contributors directly shape how state-of-the-art models perform and stay on the cutting edge of AI safety and product development.
Work is 100% remote and often flexible, letting you choose hours and workload.
Many projects require no prior experience; specialized roles pay more for domain expertise.
Your contributions help make AI systems safer, fairer, and more robust.
The role
OpenTrain is recruiting an AI Red Teaming Expert to perform adversarial testing of conversational AI systems. This is text-based work focused on probing models with jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation to surface vulnerabilities automated checks miss.
You will produce reproducible attack cases, datasets, and reports that downstream teams can act on. This is a contractor, part-time role handled through OpenTrain and is open to contributors worldwide.
Engagement type: Contractor, Part-time
Pay: $17–$25 USD per hour
Default commitment: 40 hours per week (minimum expectation: 20+ hours/week)
Work type: Remote, text-based adversarial testing
Languages required: Native fluency in English and Vietnamese
Data: Text labeling/evaluation (RLHF, evaluation rating)
What you'll do
You will run structured adversarial evaluations of conversational agents, document failures, and create reusable artifacts teams can use to fix issues.
Design and execute adversarial prompts, jailbreaks, and multi-turn attacks against conversational models.
Annotate failures, classify vulnerability types, and flag systemic risks in model outputs.
Follow taxonomies, benchmarks, and testing playbooks to keep evaluations consistent and reproducible.
Create clear, reproducible reports, datasets, and attack cases for downstream remediation and training.
Provide concise explanations of risks for both technical and non-technical stakeholders.
Requirements
You must bring hands-on adversarial testing experience plus the ability to work consistently within structured frameworks and communicate findings clearly.
Prior experience in AI red teaming, adversarial testing, cybersecurity, penetration testing, or socio-technical probing (required).
Comfort working on sensitive, text-based AI safety content and adversarial scenarios.
Experience using frameworks, benchmarks, taxonomies, or playbooks to structure evaluations.
Strong written communication for reproducible reports and stakeholder-facing summaries.
Native fluency in English and Vietnamese (required).
Who should apply
Apply if you enjoy adversarial thinking, can find subtle failure modes in multi-turn conversations, and can turn exploratory testing into structured, repeatable artifacts. This role suits experienced red teamers, security researchers, and machine-learning practitioners who prefer text-based AI safety work and who are fluent in both English and Vietnamese.
Experienced red teamers or security researchers with conversational AI experience.
Machine-learning practitioners comfortable with model behavior, emergent failure modes, and RLHF-style evaluation.
People who produce clear documentation and reproducible datasets from exploratory testing.
How it works
Create a free OpenTrain account, build your profile to reflect your red teaming and adversarial testing experience, and apply through the platform. Work assignments are managed through OpenTrain; you will submit annotations, ratings, and reports as specified by the project's playbooks and deliverables.
Compensation is hourly at the stated rate and handled via OpenTrain for contractor engagements. Your OpenTrain profile will accumulate proof-of-work from completed projects to help you find future roles in AI training and data labeling.
Join OpenTrain to red team conversational AI—find jailbreaks, prompt injections, and multi-turn manipulation, and produce reproducible vulnerability reports; remote, contract, 20+ hrs/week, $48–$62/hr. Native English and Danish required.
Join OpenTrain as an AI Safety Red Team Expert to probe conversational AI with adversarial prompts, jailbreaks, and structured red‑team methods. Remote, contract, 20+ hrs/week, pay $29–$45/hr; native English and Portuguese required.
Join OpenTrain AI as an expert red teamer probing conversational models for jailbreaks, prompt injection, bias exploitation, and multi-turn manipulation; remote contractor role, 20+ hrs/week, $48–$62/hr, native English and Dutch required.