Join us for a talk by Xuandong Zhao from UC Berkeley on the latest research into safety alignment and robustness in large language models (LLMs).

Abstract: Safety alignment in large language models (LLMs) often appears robust, but recent evidence suggests that this robustness is only shallow. In this talk, I will discuss how seemingly well-aligned models can be systematically jailbroken and how new training methods can address these vulnerabilities. I’ll begin with the Weak-to-Strong Jailbreaking study, which shows that small “unsafe” models can manipulate the decoding process of much larger aligned models to generate harmful outputs, revealing that current alignment primarily acts on initial token distributions rather than deep generative behavior. I’ll then present Dual-Objective Optimization for Refusal (DOOR), a new alignment framework that combines robust refusal learning and targeted unlearning to improve safety depth. Together, these works expose the limits of surface-level alignment and offer a pathway toward more resilient, token-level safety mechanisms in next-generation LLMs.

Speaker Bio: Xuandong Zhao is a Postdoctoral Researcher at UC Berkeley, affiliated with the Center for Responsible Decentralized Intelligence (RDI) and the Berkeley Artificial Intelligence Research (BAIR) Lab, where he works with Prof. Dawn Song. He earned his Ph.D. in Computer Science from UC Santa Barbara, where he was advised by Prof. Yu-Xiang Wang and Prof. Lei Li. His research lies at the intersection of machine learning, natural language processing, and AI safety, with a particular emphasis on responsible and reliable generative AI. Xuandong has published over 40 papers in top-tier venues spanning machine learning, security, and natural language processing. He has served as an Area Chair for conferences such as ACL. He is a recipient of the Chancellor’s Fellowship from UCSB and has been recognized with Rising Star awards in adversarial machine learning and AI.