AI Safety Tester / Prompt Analyst (Adversarial Robustness Red Teaming Report)
Conducted an adversarial robustness red-teaming exercise against a frontier LLM using only natural-language semantic structures. Designed multi-phase prompt attacks to probe safety/guardrail behavior, including safety reverse-engineering, moral/context framing, context switching toward sensitive topics, and emergency simulation to test priority overrides. Documented and reported outcomes of each phase, with an external audit evaluating the candidate’s lateral thinking, vulnerability understanding, and role suitability for AI safety testing. • Meta-prompting to elicit the model’s operational safety boundaries • Context framing using moral paradox scenarios to bypass contextual defenses • Trojan-horse style context switching requests for sensitive information • Social engineering via emergency simulation to trigger maximum-priority behavior