Probe conversational AI for jailbreaks, prompt injections, bias exploitation, and harmful behavior. Use structured evaluations and reproducible reports to help strengthen AI safety, working 20+ hours per week at $16 to $22 per hour.
About OpenTrain
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes.
Creating an OpenTrain account is free, and your profile can help you develop a lasting portfolio of experience in the rapidly growing AI training industry.
About AI Safety Training
AI training is the human side of building modern artificial intelligence. People review model behavior, create evaluation data, and identify failures that automated systems may not catch.
In this role, your safety-focused feedback will help expand how conversational AI systems are tested for harmful, biased, misleading, or manipulative behavior.
- Remote contractor work with flexible scheduling
- Part-time commitment of 20+ hours per week
- Direct involvement in cutting-edge conversational AI evaluation
The Role
OpenTrain AI is seeking an AI Safety Red Team Expert to probe conversational AI models and agents with adversarial inputs. You will test how systems respond to jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
You will review text-based outputs involving sensitive topics such as bias, misinformation, and harmful behavior. Your work will generate structured human data and reproducible artifacts that help strengthen AI systems and broaden their safety evaluation coverage.
- Role type: Part-time contractor
- Experience level: Entry level
- Pay: $16 to $22 USD per hour
- Data type: Text
What You'll Do
You will apply evaluation taxonomies, benchmarks, playbooks, and quality standards consistently across projects and task types. Your findings should be clear enough for both technical and non-technical audiences to understand and act on.
- Probe conversational AI models and agents for jailbreaks and prompt injections
- Test misuse cases, bias exploitation, and multi-turn manipulation
- Annotate model failures and classify vulnerabilities
- Identify systemic risks, subtle errors, inconsistencies, and gaps
- Review whether responses are accurate, complete, appropriate, or harmful
- Produce reproducible reports, datasets, and attack cases
- Identify vulnerabilities that automated tests may miss
- Help broaden scenario coverage across evaluation tasks
Requirements
This is an entry-level opportunity requiring strong judgment, careful reasoning, and the ability to follow structured guidance. Native fluency in both English and Assamese is required for nuanced language and content review.
- Native fluency in English and Assamese
- Strong judgment when assessing AI responses for accuracy, completeness, appropriateness, and harm
- Ability to identify and classify vulnerabilities, subtle errors, inconsistencies, and systemic risks
- Consistent application of taxonomies, benchmarks, playbooks, and quality standards
- Clear communication with technical and non-technical audiences
- Adaptability across projects, task types, and evaluation scenarios
Helpful Backgrounds
Relevant experience may include adversarial machine learning, jailbreak datasets, prompt injection, model extraction, cybersecurity, penetration testing, exploit development, reverse engineering, socio-technical risk, misinformation probing, abuse analysis, or conversational AI testing.
Backgrounds in psychology, acting, or writing may also support the unconventional adversarial thinking needed to explore how models can be manipulated.
- Adversarial machine learning or conversational AI testing
- Cybersecurity, penetration testing, exploit development, or reverse engineering
- Prompt injection, jailbreak datasets, or model extraction
- Socio-technical risk, misinformation probing, or abuse analysis
- Psychology, acting, or writing
Eligibility and Work Details
This contractor opportunity is available to applicants located in the following countries: Austria, Belgium, Bulgaria, Canada, Cyprus, Czech Republic, Germany, Denmark, Estonia, Spain, Finland, France, United Kingdom, Greece, Croatia, Hungary, Ireland, Italy, Lithuania, Luxembourg, Latvia, Malta, Netherlands, Poland, Portugal, Romania, Sweden, Slovenia, Slovakia, and the United States.
- Minimum time requirement: 20+ hours per week
- Hourly compensation: $16 to $22 USD
- Languages used: English and Assamese
- Contract arrangement: Part time
Build Your AI Training Career
AI training and data-labeling work is one of the fastest-growing ways to participate in technology. Contributors help shape how state-of-the-art systems understand information, respond to people, and handle difficult real-world situations.
Through OpenTrain, you can build a profile around your AI training experience, discover projects aligned with your skills, and apply to opportunities in minutes.
- Create a free OpenTrain account
- Showcase your evaluation and safety experience
- Apply for AI training opportunities through OpenTrain