Join OpenTrain as an AI Safety Red Team Specialist to adversarially test conversational models, document reproducible jailbreaks and misuse cases, and help make AI safer. Part-time contract, 20+ hrs/week, $20–$22/hr; native fluency in English and Assamese required.
Generative AI & RLHF
100% Remote Hourly · $20–$22/hr
$20–$22/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 13, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for building careers in AI training and data labeling. We connect skilled contributors with applied AI-safety work, help people track and grow their portfolios, and make it easy to take on meaningful, paid projects in a single place.
We hire and contract directly for this role. Working with OpenTrain means joining a fast-growing community that helps shape how modern AI systems are tested and improved.
About AI training and red teaming
AI training (data labeling, annotation, and human-feedback work) is the human side of building AI: people craft examples, probe model behavior, and evaluate outputs so models learn to be safer and more reliable.
Red teaming is adversarial testing focused on finding vulnerabilities—jailbreaks, prompt injections, bias exploits and multi-turn manipulation—so developers can fix weaknesses before they reach users.
The role: AI Safety Red Team Specialist
You will adversarially test conversational models and agents using written prompts and multi-turn attacks, classify failure modes, and produce reproducible attack cases, datasets, and reports that help engineers mitigate risk.
This is a part-time contractor role (20+ hours/week) that is fully remote and worldwide. Work is text-based and centers on clear, careful documentation suitable for technical and non-technical audiences.
Employment type: Contractor, Part-time
Time requirement: 20+ hours per week
Pay: $20–$22 USD per hour (PAY_PER_HOUR)
Data type: Text; label types include RLHF, Evaluation Rating, Red Teaming
What you'll do day-to-day
Follow structured taxonomies, benchmarks, and playbooks to keep testing consistent and reproducible. Produce clear reports and labeled examples that capture how and why attacks succeed.
Red team conversational AI with adversarial prompts and multi-turn attacks
Annotate failures and classify vulnerabilities in model outputs
Flag systemic risks like bias, misinformation, or harmful behaviors
Document reproducible attack cases and produce datasets and reports
Work on sensitive topics under explicit guidance and safety rules
Requirements
You must preserve the following skills and attributes from the role description; we will verify language fluency and red-teaming experience during onboarding.
Prior experience in AI red teaming, adversarial testing, cybersecurity, or socio-technical probing
Ability to probe jailbreaks, prompt injections, and model misuse cases
Comfort working with structured frameworks, benchmarks, and playbooks
Strong written judgment: explain risks clearly to technical and non-technical stakeholders
Native fluency in English (en) and Assamese (as)
Helpful background (not required)
You’ll stand out if you have hands-on experience in adversarial ML or cybersecurity, or if your background includes socio-technical risk analysis.
Adversarial ML experience (jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction)
Cybersecurity skills such as penetration testing or exploit development
Socio-technical risk analysis of harassment, disinformation, or abuse
Creative probing skills from psychology, acting, writing, or similar disciplines
Who should apply
This role is a fit for careful, curious people who enjoy adversarial thinking and structured annotation. It’s suitable for entry-level contributors with prior red-teaming or related experience and for experienced practitioners seeking part-time contractor work.
OpenTrain supports contributors building long-term careers in AI training—this project is a practical way to gain experience working on high-impact AI-safety problems.
How it works: onboarding, pay, and safety
You will complete onboarding that verifies language fluency and domain knowledge and trains you on playbooks and taxonomies we use for consistent labeling.
Pay is hourly at $20–$22 USD and this contract role requires 20+ hours per week. As you contribute, you will produce labeled examples, attack cases, and written reports used to improve model safety.
Onboarding includes testing and training on playbooks and reporting standards
Work remotely on text-only tasks and follow strict safety and handling guidelines
Compensation: paid hourly; exact rate within the listed range depends on assessment
Join OpenTrain as an AI Safety Red Team Specialist testing conversational models with adversarial prompts and jailbreaks; remote, 20+ hrs/week, $20–$22/hr. Requires expert red-teaming experience and native fluency in English and Punjabi.
Join OpenTrain AI to adversarially test conversational models, surface vulnerabilities, and produce reproducible attack cases; remote contractor work paying $20–$22/hr for experienced red teamers fluent in English and Bengali.
OpenTrain is hiring an experienced Red-Teaming QA Lead to audit adversarial prompts, safety evaluations, and trainer submissions—providing precise rubric-based feedback and improving contributor consistency. Remote, US-only contract role at $100/hr, 20+ hours/week.