Test conversational AI with jailbreaks, prompt injections, and other adversarial techniques while documenting vulnerabilities that improve model safety. This remote, part-time contractor role pays $16-$22 per hour and requires native English and Urdu fluency.
About OpenTrain
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. Create a free profile to discover cutting-edge projects, showcase your experience, and apply in minutes.
- Build a lasting portfolio of AI training and safety work
- Find projects that match your skills and language experience
- Work remotely as an independent contractor
About AI Safety Training
AI training is the human side of building artificial intelligence. Contributors test systems, evaluate model outputs, and prepare structured data that helps AI become more accurate, useful, and safe.
Red teaming is a specialized form of model evaluation. Human testers deliberately probe conversational AI with adversarial scenarios to uncover weaknesses that automated checks may miss.
- Help shape how conversational AI handles challenging situations
- Contribute to safer systems through structured human feedback
- Work remotely with flexible scheduling around a 20-plus-hour weekly commitment
The Role
OpenTrain is seeking an AI Safety Red Teaming Expert to test conversational AI models and agents with adversarial inputs. You will identify weaknesses, classify vulnerabilities, and create structured human data that supports safer AI systems.
This text-based contractor role is listed as entry level and is available part time at $16 to $22 per hour. Higher-sensitivity assignments are optional and include clear topic guidance and wellness resources.
- Role type: Part-time contractor
- Experience level: Entry level
- Time commitment: 20 or more hours per week
- Pay: $16-$22 per hour
- Data type: Text
- Work format: Remote
What You'll Do
You will conduct structured adversarial testing and turn your findings into clear, reproducible documentation. The work includes evaluating model behavior across safety, accuracy, appropriateness, and completeness criteria.
- Test conversational AI models and agents with jailbreaks and prompt injections
- Probe misuse cases, bias exploitation, sensitive topics, and multi-turn manipulation
- Annotate model failures and classify vulnerabilities
- Flag systemic risks and identify gaps in evaluation coverage
- Apply taxonomies, benchmarks, playbooks, and quality standards consistently
- Document reproducible reports, datasets, and attack cases
- Identify vulnerabilities that automated testing may miss
- Collect and structure human data for AI safety improvements
Requirements
Native fluency in both English and Urdu is required. You should be able to make careful judgments about AI responses and explain your reasoning clearly to technical and non-technical audiences.
The role requires consistent, structured work against defined guidelines and quality standards. You should be comfortable noticing subtle errors, inconsistencies, unsafe behavior, and incomplete or inappropriate responses.
- Native fluency in English and Urdu
- Strong judgment when evaluating AI responses for accuracy, completeness, appropriateness, and safety
- Ability to identify jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Consistent use of taxonomies, benchmarks, playbooks, and quality standards
- Ability to explain vulnerability assessments and reproducible attack cases clearly
- Careful attention to subtle errors and inconsistencies
Helpful Backgrounds
Relevant experience can help you approach model testing from different angles, although these backgrounds are described as helpful rather than required. Unconventional adversarial thinking and strong written reasoning are valuable in this work.
- Adversarial machine learning, including jailbreak datasets or prompt injection
- RLHF or DPO attacks
- Model extraction
- Cybersecurity, penetration testing, exploit development, or reverse engineering
- Socio-technical risk or abuse analysis
- Harassment or misinformation probing
- Conversational AI testing
- Psychology, acting, or writing
Location And Language Eligibility
This project requires English and Urdu fluency and is available to contractors in the following countries.
- Austria, Belgium, Bulgaria, Canada, Cyprus, Czechia, Germany, Denmark, Estonia, Spain, Finland, France, United Kingdom, Greece, Croatia, Hungary, Ireland, Italy, Lithuania, Luxembourg, Latvia, Malta, Netherlands, Poland, Portugal, Romania, Sweden, Slovenia, Slovakia, and the United States
Why Build Your AI Training Career With OpenTrain
AI safety work sits at the forefront of a rapidly growing technology industry. Through OpenTrain, you can develop a profile that shows credible experience, find projects aligned with your abilities, and build a career portfolio in AI training and data labeling.
- Create an OpenTrain account for free
- Apply to projects in minutes
- Showcase your AI safety and evaluation experience
- Grow from individual projects toward a durable AI training portfolio