Test conversational AI for jailbreaks, prompt injections, bias, misinformation, and harmful behavior as a remote English and Marathi red teamer earning $16-$22 per hour.
About OpenTrain
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help contributors discover cutting-edge projects, build a professional AI training profile, and apply in minutes. Creating an OpenTrain account is free.
- Build a portfolio around specialized AI safety and evaluation work
- Find opportunities that match your language and technical strengths
- Grow experience in a rapidly expanding field shaping how AI systems behave
About AI Safety Red Teaming
AI training is the human side of building artificial intelligence. Red teamers challenge conversational models with realistic and adversarial inputs, then document how systems respond so teams can identify weaknesses, reduce harmful behavior, and improve safety.
- Work remotely with text-based AI training tasks
- Evaluate model responses for accuracy, completeness, appropriateness, bias, and safety
- Contribute structured human feedback that supports stronger AI systems
The Role
OpenTrain is seeking an English and Marathi AI Safety Red Teamer to probe conversational AI models and agents. You will test jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation while producing structured data that helps identify vulnerabilities and systemic risks.
This is remote, text-based freelance work. The default commitment is 40 hours per week, with a stated time requirement of 20 or more hours per week. Pay ranges from $16 to $22 per hour. Participation in higher-sensitivity content is optional and supported by clear topic guidance and wellness resources.
- Remote contractor role
- Part-time engagement with 20 or more hours per week
- Default commitment of 40 hours per week
- $16-$22 per hour
- English and Marathi language requirements
What You'll Do
You will follow established taxonomies, benchmarks, and playbooks to keep testing consistent across projects and task types. Your work will include reviewing model behavior, classifying failures, and documenting reproducible findings for AI safety evaluation.
- Test conversational AI models and agents with adversarial prompts
- Attempt jailbreaks, prompt injections, misuse cases, and multi-turn manipulation
- Review outputs for accuracy, completeness, appropriateness, bias, misinformation, and harmful behavior
- Annotate failures and classify vulnerabilities and systemic risks
- Expand evaluation coverage across relevant safety concerns
- Create reproducible reports, datasets, and attack cases
- Apply taxonomies, benchmarks, playbooks, guidelines, and quality standards consistently
- Explain adversarial testing decisions clearly to technical and non-technical audiences
Required Qualifications
You should be natively fluent in both English and Marathi and have strong judgment about language and content quality. The role requires careful evaluation of conversational AI responses, including subtle errors, inconsistencies, omissions, bias, and safety risks.
- Native fluency in English and Marathi
- Ability to identify jailbreaks, prompt injections, misuse cases, and multi-turn manipulation
- Strong judgment when assessing whether AI responses are accurate, complete, appropriate, and safe
- Ability to classify vulnerabilities and systemic risks using defined frameworks
- Clear written communication for technical and non-technical audiences
- Ability to apply guidelines and quality standards consistently
- Adaptability across projects and task types
Helpful Background
The role is suitable for entry-level applicants who meet the core language and judgment requirements. Additional experience can help you approach adversarial testing from different perspectives, but the following backgrounds are described as helpful rather than required.
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction
- Cybersecurity, including penetration testing, exploit development, or reverse engineering
- Socio-technical risk, harassment, misinformation, or abuse analysis
- Conversational AI testing
- Psychology, acting, or creative writing that supports unconventional adversarial thinking
Why This Work Matters
Modern AI systems depend on people who can prepare examples, evaluate outputs, and identify where models fail. By testing conversational AI in English and Marathi, you will help reveal vulnerabilities and safety risks that may otherwise be missed.
- Work directly on how state-of-the-art AI systems are evaluated
- Use language expertise to surface culturally and linguistically specific risks
- Develop experience in AI safety, red teaming, and human feedback work
- Choose flexible work that can fit around other commitments
How to Get Started
Create a free OpenTrain account, build your profile around your English and Marathi fluency and relevant experience, and apply through OpenTrain. Your profile can help you show credible AI training experience and discover future projects aligned with your skills.
- Create your free OpenTrain account
- Highlight bilingual fluency and relevant safety, cybersecurity, or evaluation experience
- Apply in minutes
- Build a long-term portfolio in AI training and data labeling