Research position with Prof. Han Liu, MAGICS Lab @ Northwestern University (Evo-MARL/AdvEvo-MARL safety co-evolution)
Conducted research to improve internalized safety in LLM-based multi-agent systems by co-training attackers and defenders in a multi-agent reinforcement learning setup. Implemented co-evolutionary training components that shape model behavior for safety, task performance, and response formatting. Evaluated adversarial scenarios to quantify attack success rates and measure downstream task utility. • Co-developed Evo-MARL and AdvEvo-MARL training frameworks for internalized safety. • Designed multi-objective reward shaping across safety, task utility, and response formatting. • Built baseline mechanisms for variance reduction via shared mean-return baselines among agents. • Reported evaluation results showing reduced attack success and improved benchmark performance.