Join OpenTrain as an AI Safety Red Team Expert to probe conversational AI with adversarial prompts, jailbreaks, and structured red‑team methods. Remote, contract, 20+ hrs/week, pay $29–$45/hr; native English and Portuguese required.
Generative AI & RLHF
100% Remote Hourly · $29–$45/hr
$29–$45/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Jul 30, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We hire and contract contributors directly and help people grow durable freelance careers in the human side of AI development.
Creating an OpenTrain account is free. We centralize projects, let you build a portfolio that follows your work, and connect you to flexible, remote opportunities across the fast-growing AI training industry.
About AI training and red teaming
AI training (data labeling, annotation, and human feedback) is how people teach models to behave. Red teaming is a specialized subset focused on adversarial testing: finding jailbreaks, prompt injections, bias exploits, and multi-turn manipulations so models can be made safer.
This role places you at the cutting edge of model safety—your findings help shape how AI systems handle sensitive topics like misinformation, bias, and harmful behavior.
The role
OpenTrain is hiring an AI Safety Red Team Expert to test conversational AI models and agents using adversarial prompts, jailbreaks, prompt injection techniques, misuse-case scenarios, and multi-turn attacks. You will identify failures, classify vulnerabilities, and produce reproducible reports and datasets.
This is a remote, contract, part-time role requiring 20+ hours per week. Pay is hourly, USD $29–$45/hr (top rate $45/hr). You must be a native speaker of English and Portuguese.
What you'll do
Probe conversational systems with adversarial inputs and structured red‑team methods to reveal weaknesses and edge-case behavior.
Annotate model failures, classify vulnerabilities, and flag systemic risks for follow-up engineering or policy work.
Follow taxonomies, benchmarks, and pre-defined playbooks to keep evaluations consistent and reproducible.
Produce clear, reproducible reports, datasets, and attack cases that developers and safety teams can act on.
Work through text-based safety review tasks that may include sensitive topics such as bias, misinformation, and harmful behavior.
Requirements
Prior red teaming experience with AI models, agents, adversarial systems, or relevant cybersecurity research.
Familiarity with jailbreaks, prompt injection, misuse cases, bias exploitation, or similar adversarial techniques.
Comfort using structured evaluation methods: taxonomies, benchmarks, playbooks or equivalent frameworks.
Clear written communication for documenting vulnerabilities and explaining systemic risk to technical and non-technical stakeholders.
Native fluency in English and Portuguese.
Availability for 20+ hours per week; this is an intermediate-level, contractor role.
Who should apply
Apply if you have hands-on experience testing models, a background in security research or ML safety, or if you have practical experience probing socio-technical systems for misuse. This role suits people who enjoy methodical adversarial testing and translating results into actionable remediation.
You do not need to be an engineer to contribute, but prior red‑teaming or adversarial testing experience is required. If you can adapt across projects, document reproducible attack cases, and communicate risk clearly, we want to hear from you.
How it works
OpenTrain hires and contracts contributors directly. You will complete text-based tasks via the OpenTrain platform, follow provided playbooks and taxonomies, and submit reproducible artifacts (reports, labeled datasets, attack cases).
To apply, create a free OpenTrain account, complete your profile, and submit your application. Work is remote and flexible within the 20+ hours/week expectation; specifics of task cadence and deadlines vary by project.
Labeling/data type: Text (RLHF, evaluation rating, text generation tasks).
Employment: Contractor, part-time. Payment: hourly in USD, $29–$45/hr.
Worldwide applicants accepted; native English and Portuguese required.
Join OpenTrain as an AI Safety Red Team Specialist testing conversational models with adversarial prompts and jailbreaks; remote, 20+ hrs/week, $20–$22/hr. Requires expert red-teaming experience and native fluency in English and Punjabi.
Join OpenTrain AI to red-team conversational models—finding jailbreaks, prompt injections, bias, and misuse—while producing reproducible attack cases and risk reports. Contract, part-time role (20+ hrs/week) paying USD 17–25/hr; native English and Indonesian required.
Join OpenTrain AI as an expert red teamer probing conversational models for jailbreaks, prompt injection, bias exploitation, and multi-turn manipulation; remote contractor role, 20+ hrs/week, $48–$62/hr, native English and Dutch required.