AI Safety & Red Teaming Project (as part of AI training/evaluation work)
Participated in an AI safety and red teaming effort to identify vulnerabilities in prototype conversational models. Attempted to bypass or break safety filters to expose risks involving harmful content, privacy leaks, and misinformation. Used findings to improve safety-related evaluation and training coverage. • Tested safety filter bypass attempts • Assessed exposure of privacy leakage risks • Evaluated misinformation and harmful content vulnerabilities • Reported vulnerability themes for mitigation