AI Safety Red Team Specialist (Conversational Models)
Join OpenTrain as an AI Safety Red Team Specialist testing conversational models with adversarial prompts and jailbreaks; remote, 20+ hrs/week, $20–$22/hr. Requires expert red-teaming experience and native fluency in English and Punjabi.
Generative AI & RLHF
100% Remote Hourly · $20–$22/hr
$20–$22/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Jul 13, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people building careers in AI training and data labeling. We connect independent contributors with specialized projects that directly shape how AI systems behave, while helping freelancers build a durable portfolio of evaluation work.
About AI training and red teaming
AI training (also called data labeling or human feedback work) is the human layer behind modern models — people prepare examples, evaluate outputs, and flag failures that automated tests miss. Red teaming is a specialized strand of that work: adversarially probing models to find jailbreaks, prompt-injection paths, bias, misinformation, and multi-turn manipulation.
This role focuses on text-based adversarial testing of conversational AI and producing reproducible attack cases, reports, and datasets that help safety teams fix vulnerabilities.
The role
You will work as an expert AI Safety Red Team Specialist running structured adversarial evaluations against conversational models and agents. The contract is part-time (20+ hours per week), remote, and open worldwide. You will document vulnerabilities, classify failure modes, and create reproducible materials for safety review.
Employment type: Contractor, Part-time
Time requirement: 20+ hours/week
Data type: Text; label types: RLHF, Evaluation Rating
Pay: $20–$22 USD per hour
What you'll do
Your day-to-day work is text-focused adversarial testing and clear documentation so engineering and safety teams can reproduce and remediate issues.
Design and execute adversarial prompts, jailbreaks, prompt-injection and multi-turn attack sequences following playbooks and taxonomies
Annotate failures and classify vulnerabilities using existing frameworks and benchmarks
Flag systemic risks and sensitive behavior (bias, misinformation, harmful instructions) while following safety guidelines
Produce reproducible reports, labeled datasets, and attack cases for model safety review and tracking
Requirements
We require demonstrable, expert-level experience and language skills. Preserve and reflect the structured testing habits you bring to every evaluation.
Prior red teaming or adversarial model-evaluation experience (AI adversarial work, cybersecurity, or socio-technical probing)
Familiarity with jailbreaks, prompt injections, misuse cases, and bias exploitation
Structured use of frameworks, taxonomies, benchmarks, or playbooks to ensure consistent testing
Clear written communication for technical and non-technical stakeholders
Native fluency in English (en) and Punjabi (pa)
Helpful background (not required but valued)
The following backgrounds will make you especially effective at this role and help you design creative, high-impact adversarial tests.
Adversarial ML experience (jailbreak datasets, prompt-injection research, RLHF/DPO attacks, model extraction)
Experience probing harassment, disinformation, abuse, or other conversational AI risks
Creative adversarial thinking from psychology, acting, writing, or social engineering
How the project works, compensation, and next steps
OpenTrain hires contractors for part-time schedules and pays on an hourly basis. Work is performed remotely and may touch sensitive topics; you will follow provided safety and content guidelines when annotating and reporting. If you meet the requirements, apply with examples of prior red-team or adversarial evaluation work and a short note about how you structure tests and document reproducibility.
Compensation: $20–$22 USD per hour, paid per contractor agreement
Worldwide applicants accepted; you must be native-fluent in English and Punjabi
Labeling context: text-based RLHF and evaluation rating tasks; expect detailed reporting and dataset creation
Join OpenTrain AI to adversarially test conversational models, surface vulnerabilities, and produce reproducible attack cases; remote contractor work paying $20–$22/hr for experienced red teamers fluent in English and Bengali.
Join OpenTrain as an AI Safety Red Team Specialist to adversarially test conversational models, document reproducible jailbreaks and misuse cases, and help make AI safer. Part-time contract, 20+ hrs/week, $20–$22/hr; native fluency in English and Assamese required.
Join OpenTrain as an AI Safety Red Teamer probing conversational models for jailbreaks, prompt injections, and misuse. Remote contractor role, 20+ hrs/week, $20–22/hr; native fluency in English and Odia required.