AI Safety & Red-Teaming Specialist | DataAnnotation.Tech (Remote)
Conducted adversarial testing to generate and validate jailbreak prompts that attempt to bypass model safety filters and operational guidelines. Documented vulnerability vectors and provided actionable feedback to patch alignment flaws. Performed exhaustive fact-checking on technical model outputs to minimize hallucinations. • Built sophisticated jailbreak prompts for safety filter circumvention testing • Recorded vulnerability patterns such as roleplay circumvention, payload splitting, and base64 encoding attacks • Fact-checked outputs using robust research methodologies in technical domains • Supplied engineering-ready guidance to remediate alignment vulnerabilities