Probe conversational AI models for jailbreaks, prompt injections, bias, and other vulnerabilities while creating data that improves AI safety. This part-time contractor role requires native English and Tamil fluency and pays $16 to $22 per hour.
About OpenTrain
OpenTrain AI is the hiring and contracting organization for specialized AI training and data-labeling work. It helps contributors build experience in a fast-growing field, organize their AI training careers, and apply their skills to projects that shape how modern AI systems behave.
- Build experience in human data-driven AI safety and evaluation.
- Create an OpenTrain account for free and apply in minutes.
- Work as a part-time contractor on specialized AI training assignments.
About AI Training and Red Teaming
AI training is the human side of building artificial intelligence. People evaluate model responses, identify failures, classify examples, and provide structured feedback so AI systems become more accurate, reliable, and safe.
Red teaming applies adversarial thinking to this process. By testing conversational models with challenging inputs and documenting reproducible failures, contributors help reveal weaknesses that automated tests may not detect.
- Evaluate text-based conversational AI outputs.
- Help improve the safety and robustness of emerging AI systems.
- Contribute to cutting-edge work with flexible part-time availability.
The Role
OpenTrain is recruiting a Conversational AI Red Teaming Expert to probe conversational AI models and agents with adversarial inputs, uncover vulnerabilities, and create human-generated data that improves AI safety. The work combines adversarial testing, structured evaluation, vulnerability classification, and reproducible documentation.
You will review text-based model outputs involving jailbreaks, prompt injections, misuse cases, bias exploitation, misinformation, harmful behaviors, and multi-turn manipulation. The role is listed as entry level and requires at least 20 hours per week.
- Contractor, part-time position
- Pay: $16 to $22 USD per hour
- Data type: Text
- Available in the listed countries across Europe, Canada, and the United States
What You’ll Do
You will apply structured testing methods and quality standards to uncover, assess, and communicate conversational AI weaknesses. Your work will support broader evaluation coverage and help reduce unexpected failures in production systems.
- Probe models and agents for jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
- Review responses for accuracy, completeness, appropriateness, and systemic risk.
- Annotate failures and classify vulnerabilities using defined taxonomies.
- Identify weaknesses that automated tests may miss.
- Apply benchmarks, playbooks, and quality standards consistently across scenarios.
- Produce reproducible attack cases, datasets, and clear reports.
- Expand evaluation coverage for conversational AI systems.
Requirements
Native fluency in both English and Tamil is required, along with strong written communication. You should be able to assess language, content quality, safety, and model behavior with careful attention to subtle errors, inconsistencies, and gaps.
You must be comfortable following structured guidelines and explaining analytical decisions clearly to both technical and non-technical audiences.
- Native English and Tamil fluency
- Strong judgment when evaluating conversational AI responses
- Ability to assess accuracy, completeness, appropriateness, and safety
- Ability to recognize jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Ability to apply structured taxonomies and quality standards consistently
- Clear, rigorous reporting of vulnerability findings
Helpful Background
Experience in adversarial machine learning, jailbreak datasets, prompt injection, RLHF or DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, abuse analysis, or conversational AI testing is useful but not required.
Backgrounds in psychology, acting, or creative writing may also support the unconventional adversarial thinking needed for this work. Some assignments may involve sensitive topics such as bias, misinformation, harassment, or harmful behavior; the role includes clear topic guidance and wellness resources.
- Adversarial machine learning or conversational AI testing
- Jailbreak datasets or prompt injection
- RLHF or DPO attacks
- Penetration testing, exploit development, or reverse engineering
- Abuse analysis
- Psychology, acting, or creative writing
Why Work With OpenTrain
This role offers a way to build practical experience in human data-driven AI red teaming while contributing directly to safer, more robust, and more trustworthy AI systems. AI training work can be a flexible path into technology, with opportunities to develop specialized skills through projects that influence state-of-the-art models.
- Contribute to safer and more trustworthy conversational AI.
- Develop specialized experience in AI red teaming and evaluation.
- Work part time for 20 or more hours per week.
- Grow a durable freelance career in AI training.
How to Apply
Create a free OpenTrain account, build your profile around your language and evaluation experience, and apply for this contractor opportunity in minutes. Be prepared to demonstrate native English and Tamil fluency and your ability to evaluate conversational AI behavior against structured safety and quality standards.
- Review the language and availability requirements.
- Create or update your OpenTrain profile.
- Apply through OpenTrain for consideration.