Corpus Construction & AI Safety Specialist | Remote (AI red-teaming and safety data work)
Performed adversarial testing and safety evaluation by generating and operationalizing attack prompts against LLMs. Documented and categorized vulnerabilities related to regional bias, NSFW content exposure, and privacy leakage to inform mitigation. Hardened model safety boundaries by developing optimized refusal-response behaviors for regulated generative AI scenarios. • Built and tested 200+ high-level attack prompts using logic sandboxing and dialect-nested slang • Conducted jailbreak detection and refusal-response hardening for compliance • Identified vulnerability categories including bias, NSFW, and privacy leaks • Ensured red-teaming activities were aligned with generative AI regulations