Join OpenTrain as an AI Safety Red Teamer testing conversational models with jailbreaks, prompt injections, and adversarial attacks. This contract, remote role pays $20–$22/hr, requires native English and Odia, and expects 20+ hours/week.
Generative AI & RLHF
100% Remote Hourly · $20–$22/hr
$20–$22/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Jul 13, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people building careers in AI training and data labeling. We connect contractors with specialized tasks that shape how real-world AI systems behave and give contributors a place to grow a lasting freelance portfolio.
For this role, OpenTrain AI is the hiring and contracting organization. You will work remotely for the platform on safety-focused red teaming assignments that produce reproducible attack cases, labeled datasets, and structured vulnerability reports.
Why AI Training and Red Teaming Matters
AI training (data labeling and human feedback) is the human side of building modern models — people create examples, evaluate outputs, and teach systems to be safer and more useful. Red teaming is a specialized subset that intentionally probes and breaks models so developers can fix weaknesses before harm occurs.
This work is remote, flexible, and directly influences model safety. Many projects require creativity, technical rigor, and careful documentation rather than formal degrees; domain expertise increases impact and pay.
The Role
You will be an AI Safety Red Teamer focused on testing conversational AI and agents for jailbreaks, prompt injection, model extraction, bias exploitation, and multi-turn manipulation. The role centers on finding reproducible failures, classifying vulnerabilities against taxonomies and benchmarks, and producing clear reports and datasets that help improve model safety.
This is a contractor, part-time role with a time expectation of 20+ hours per week and pay between $20 and $22 per hour (USD). Work is fully remote and open worldwide.
What You'll Do
Apply structured red team playbooks and adversarial frameworks to probe conversational models.
Annotate failures and classify vulnerabilities using consistent taxonomies and benchmarks so results are reproducible and comparable.
Document attack cases, produce clear write-ups and datasets, and explain findings to both technical and non-technical stakeholders.
Perform sensitive text-based output review tasks following precise guidance and safety instructions.
Probe systems with adversarial inputs, jailbreak techniques, and prompt-injection strategies.
Create reproducible examples and attach metadata for each failure mode.
Label and rate model responses using defined evaluation criteria.
Requirements
You must have prior red teaming or adversarial AI testing experience and be able to work with frameworks or benchmarks rather than ad hoc methods.
Native-level fluency in both English and Odia is required because assignments and reports will be written and evaluated in those languages.
Prior red teaming or adversarial AI testing experience (required).
Native fluency in English and Odia (required).
Strong written communication for technical and non-technical audiences (required).
Ability to apply frameworks, taxonomies, and benchmarks consistently (required).
Background in cybersecurity, penetration testing, exploit development, reverse engineering, or socio-technical risk probing is a strong plus.
Helpful Background
Experience with jailbreak datasets, prompt injection research, RLHF/DPO attack techniques, model extraction, or probing harassment/misinformation behaviors will make you especially effective.
Comfort adapting quickly to new instructions, tools, and workflows is important; you will often iterate on test cases and refine documentation to meet project standards.
How It Works / Apply
If selected, you'll receive structured guidelines, playbooks, and scoring rubrics for each assignment. Deliverables typically include labeled examples, a classification of the vulnerability, and a written report or ticket describing reproducible steps and recommended mitigations.
OpenTrain manages contracts and payments. This posting is for a contractor, part-time engagement at $20–$22/hr, 20+ hours/week, with work completed remotely and submitted through OpenTrain's workflows.
Employment: Contractor, Part-time.
Expected time: 20+ hours/week.
Pay: $20–$22 per hour (USD).
Data type: Text; label types include Red Teaming and Evaluation Rating.
Join OpenTrain to red team conversational AI: probe models with adversarial prompts, document reproducible attacks, and help improve safety. Part-time contractor role (20+ hrs/week), $20–22/hr, requires expert red teaming experience and fluency in English and Bengali.
Join OpenTrain to evaluate and create gold-standard Odia AI responses: remote, contract, part-time work at $15/hr for 20+ hours/week. Native or near-native Odia reading/writing and professional English for detailed feedback are required.
Join OpenTrain as a part-time AI Safety Red Teamer to design adversarial prompts and evaluate frontier models across high-risk domains; remote, 20+ hrs/week, paid $70–$84/hr. Requires 5+ years' relevant experience and strong written analytical skills.