AI Safety Red Team Specialist, Text Adversarial Testing
OpenTrain is hiring an expert AI Safety Red Team Specialist to perform text-based adversarial testing of conversational models (20+ hrs/week). Earn $48–$62/hr while finding jailbreaks, prompt injections, and systemic failure cases; native English and Finnish required.
Generative AI & RLHF
100% Remote Hourly · $48–$62/hr
$48–$62/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Jul 30, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people building careers in AI training and data labeling. We help contributors find specialist projects, consolidate experience across work, and build a reusable portfolio that demonstrates their expertise.
OpenTrain is hiring and contracting directly for this role. Creating an OpenTrain account is free and lets you apply, track work, and grow a durable freelance career in AI training.
Why AI Training and Red Team Work Matters
AI training is the human side of building intelligent systems — people prepare, test, and correct the examples models learn from. Red teaming is a high-impact niche: adversarial testing surfaces real-world failure modes that automated checks often miss.
This work is remote and flexible, often suitable as part-time or contract work. Contributors directly shape how state-of-the-art conversational AI behaves, improving safety, robustness, and user trust.
The Role
As an AI Safety Red Team Specialist you will perform text-based adversarial testing of conversational models and agents. Your output will include annotated failures, classified vulnerabilities, and reproducible attack cases that feed into model safety reviews and datasets.
This is a contractor, part-time role (20+ hours per week). Pay is hourly at $48–$62 USD per hour, paid per timesheet. Work is worldwide-remote; you must be natively fluent in English and Finnish.
Work type: Contractor, Part-time
Time commitment: 20+ hours/week
Pay: $48–$62 USD per hour
Languages required: Native English and Finnish
Data type: Text (RLHF and evaluation rating tasks)
What You'll Do
You will design, execute, and document adversarial test cases against conversational AI systems using multi-turn attacks and structured methodologies. Maintain reproducibility and consistency by applying taxonomies, benchmarks, and playbooks.
Red team conversational models with jailbreaks, prompt injections, misuse scenarios, and multi-turn manipulation.
Annotate failures and classify vulnerabilities using established taxonomies and benchmarks.
Flag systemic risks and prioritize issues by severity and exploitability.
Produce reproducible reports, labeled datasets, and attack cases for safety review and model improvement.
Requirements
You must be an experienced red teamer or safety evaluator with demonstrated ability to probe models and create high-quality, reproducible reporting. Preserve clear, structured evaluation habits and communicate findings to diverse stakeholders.
Prior red teaming experience in AI adversarial testing, cybersecurity, or socio-technical probing.
Demonstrated ability to test conversational models with jailbreaks, prompt injections, and misuse cases.
Strong written communication for clear technical and non-technical risk reporting.
Structured evaluation habits: use frameworks, taxonomies, or benchmarks rather than ad hoc testing.
Native fluency in English and Finnish.
Experience level: Expert
Who Should Apply
Apply if you are a security researcher, AI safety specialist, RLHF evaluator, or experienced red teamer who enjoys creative adversarial thinking and rigorous documentation. This role suits people who prefer remote, flexible contract work and want to help shape safer conversational AI.
Candidates with backgrounds in cybersecurity, red teaming, threat modeling, or RLHF evaluation are a strong fit.
Ideal for people who can turn adversarial tests into reproducible datasets and clear recommendations.
How It Works
Create a free OpenTrain account to apply. We'll evaluate your experience and may request examples of past red teaming or adversarial testing work. If selected, you will receive onboarding materials, taxonomies, and playbooks to ensure consistent evaluations.
Work is remote and tracked through OpenTrain. Payments follow the stated hourly range and are processed per contract terms. OpenTrain supports contributors as they grow their AI training careers.
Application: Free OpenTrain account and submission of experience/examples.
Onboarding: Playbooks, taxonomies, and evaluation guidelines provided.
Payments: Hourly at $48–$62 USD, contractor terms.
Join OpenTrain as an AI Safety Red Team Specialist testing conversational models with adversarial prompts and jailbreaks; remote, 20+ hrs/week, $20–$22/hr. Requires expert red-teaming experience and native fluency in English and Punjabi.
Probe conversational AI for jailbreaks, prompt injections, and misuse as a remote AI Safety Red Team Specialist for OpenTrain. Contractor, part-time role (20+ hrs/week) paying $24–$35/hr; native English and Thai required.
Join OpenTrain AI to adversarially test conversational models, surface vulnerabilities, and produce reproducible attack cases; remote contractor work paying $20–$22/hr for experienced red teamers fluent in English and Bengali.