Test conversational AI with jailbreaks, prompt injections, and other adversarial scenarios while documenting vulnerabilities that automated tools may miss. This remote contractor role requires native English and Odia fluency and pays $16 to $22 per hour.
About OpenTrain
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help contributors discover projects, build a professional profile, and apply to opportunities in minutes as they grow experience in this fast-moving field.
- Free account creation
- A profile to showcase AI training experience
- Remote opportunities across the growing AI industry
About AI Safety Training
AI training is the human side of building modern artificial intelligence. People evaluate model responses, identify weaknesses, create structured examples, and provide feedback that helps AI systems become more reliable, useful, and safe.
Red teaming applies adversarial thinking to this process. By probing conversational models and agents with difficult or manipulative inputs, contributors can uncover risks that automated testing may not detect.
- Work directly with conversational AI evaluation
- Help identify model weaknesses and systemic risks
- Contribute to the development of safer AI systems
The Role
OpenTrain is seeking an AI Safety Red Teaming Expert for text-based testing of conversational AI models and agents. You will use adversarial inputs to explore how systems respond to challenging, manipulative, or unsafe scenarios, then produce structured human data to support safety improvements.
Assignments may involve sensitive topics such as bias, misinformation, or harmful behavior. Higher-sensitivity work is supported by clear guidance and wellness resources, and topics are communicated before exposure.
- Remote contractor position
- Part-time contract classification
- Entry-level experience designation
- Pay of $16 to $22 per hour
What You'll Do
You will evaluate conversational AI across a broad range of adversarial scenarios. Your work will combine hands-on testing, careful annotation, vulnerability classification, and clear documentation that can be used by technical and non-technical teams.
- Test for jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Annotate model failures and classify vulnerabilities
- Identify systemic risks in conversational AI behavior
- Apply taxonomies, benchmarks, and playbooks consistently
- Create reproducible reports, datasets, and attack cases
- Find vulnerabilities that automated testing may miss
- Expand evaluation coverage across task types and projects
Requirements
Native fluency in both English and Odia is required. You should be able to assess whether AI responses are accurate, complete, and appropriate while noticing subtle errors, inconsistencies, and gaps.
The role also requires disciplined documentation and the ability to apply structured guidance across evaluation projects. Helpful experience includes adversarial machine learning, jailbreak datasets, prompt injection, RLHF or DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, abuse analysis, misinformation probing, or conversational AI testing.
Backgrounds in psychology, acting, or unconventional creative writing may also support the adversarial thinking needed for this work.
- Native English and Odia fluency
- Strong judgment when evaluating AI responses
- Careful attention to subtle errors and inconsistencies
- Ability to follow structured taxonomies and quality standards
- Clear communication with technical and non-technical audiences
- Ability to adapt across evaluation projects and task types
Schedule, Location, and Compensation
This is remote contractor work available to candidates in the listed countries below. The role has a stated availability requirement of 20 or more hours per week and a default commitment of 40 hours per week.
- Countries: Austria, Belgium, Bulgaria, Canada, Cyprus, Czech Republic, Germany, Denmark, Estonia, Spain, Finland, France, United Kingdom, Greece, Croatia, Hungary, Ireland, Italy, Lithuania, Luxembourg, Latvia, Malta, Netherlands, Poland, Portugal, Romania, Sweden, Slovenia, Slovakia, and the United
- Availability requirement: 20+ hours per week
- Default commitment: 40 hours per week
- Hourly rate: $16 to $22 USD
Build Your AI Training Career
AI safety red teaming is part of a rapidly growing industry where human judgment directly shapes how advanced AI behaves. Through OpenTrain, you can build a profile around your work, discover projects that match your capabilities, and develop a lasting career portfolio in AI training and data labeling.
- Apply your language skills to cutting-edge AI work
- Build experience in model evaluation and safety testing
- Create a stronger professional record through OpenTrain