Test conversational AI for jailbreaks, prompt injections, bias, misuse, and other safety risks. Work remotely for 20+ hours per week at $16 to $22 per hour.
About OpenTrain
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover specialized projects, build a credible portfolio, and grow their experience in a fast-moving field.
- Create a free OpenTrain account and apply in minutes.
- Build a profile that showcases your AI training and safety evaluation experience.
- Find flexible contract opportunities designed for long-term career growth.
About AI Safety Training Work
AI training is the human side of building artificial intelligence. People evaluate model responses, identify weaknesses, prepare high-quality examples, and provide feedback that helps AI systems become safer, more accurate, and more trustworthy.
Red teaming is a specialized form of AI evaluation. By testing models with adversarial prompts and documenting failures, contributors help technical teams find risks that automated testing may overlook.
- Work remotely with a computer and an internet connection.
- Contribute directly to the development and evaluation of modern AI systems.
- Use careful judgment to assess model behavior across challenging scenarios.
The Role
OpenTrain is seeking an AI Red Teaming Expert to probe conversational AI models and agents with adversarial inputs. You will uncover vulnerabilities, evaluate sensitive model behavior, classify risks, and produce human-generated data that supports AI safety improvements.
This is text-based remote work focused on jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation, and related safety concerns. Higher-sensitivity projects are optional and supported by clear guidelines and wellness resources.
- Role: AI Red Teaming Expert
- Experience level: Entry level
- Work arrangement: Remote contractor, part time
- Time commitment: 20+ hours per week
- Pay: $16 to $22 per hour
- Languages: Native English and Malayalam fluency required
- Eligible locations: Austria, Belgium, Bulgaria, Canada, Cyprus, Czech Republic, Germany, Denmark, Estonia, Spain, Finland, France, United Kingdom, Greece, Croatia, Hungary, Ireland, Italy, Lithuania, Luxembourg, Latvia, Malta, Netherlands, Poland, Portugal, Romania, Sweden, Slovenia, Slovakia, and t
What You'll Do
You will challenge conversational AI systems, review their outputs, and turn observed failures into clear, reproducible evidence. Your work will help expand evaluation coverage and give technical teams actionable information for improving safety and trustworthiness.
- Probe models and agents for jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
- Review outputs involving bias, misinformation, harmful behaviors, and related safety concerns.
- Annotate failures and classify vulnerabilities using taxonomies, benchmarks, playbooks, and quality standards.
- Identify systemic risks and vulnerabilities that automated tests may miss.
- Create reproducible reports, datasets, and attack cases that technical teams can act on.
- Expand scenario and evaluation coverage across different projects and task types.
Requirements
This role requires native fluency in both English and Malayalam, strong written communication, and the ability to make nuanced judgments about conversational AI responses. You must be able to explain your reasoning clearly to both technical and nontechnical audiences while following structured guidelines consistently.
- Native fluency in English and Malayalam.
- Strong written communication skills.
- Ability to assess responses for accuracy, completeness, appropriateness, and safety risks.
- Experience or familiarity with jailbreaks, prompt injection, misuse cases, bias analysis, or other adversarial AI testing.
- Ability to follow taxonomies and quality standards.
- Rigorous attention to subtle errors, inconsistencies, and gaps.
- Ability to document clear, reproducible findings and explain reasoning.
Helpful Background
The following experience or perspectives may help you approach adversarial testing creatively. These are helpful backgrounds rather than stated mandatory qualifications.
- Adversarial machine learning or jailbreak datasets
- Prompt injection or RLHF and DPO attacks
- Model extraction, penetration testing, exploit development, or reverse engineering
- Socio-technical risk analysis or abuse analysis
- Conversational AI testing
- Psychology, acting, or writing that supports unconventional adversarial thinking
What Success Looks Like
Success means uncovering meaningful vulnerabilities that automated testing misses and delivering attack cases, reports, and datasets that others can reproduce. Your findings will help expand scenario coverage, reduce unexpected production risks, and give AI teams clearer evidence for improving model safety.
- Discover subtle or systemic weaknesses in conversational AI behavior.
- Produce findings that are specific, well-supported, and reproducible.
- Apply evaluation standards consistently across projects.
- Help technical teams make better-informed safety improvements.
How to Apply Through OpenTrain
Create a free OpenTrain account, build your profile, and apply for this contract opportunity in minutes. OpenTrain helps AI training contributors present their experience, discover specialized work, and develop a durable portfolio in AI safety and data labeling.
- Confirm that you meet the English and Malayalam fluency requirement.
- Highlight relevant adversarial AI testing, safety evaluation, or analytical experience.
- Apply through OpenTrain for consideration.