From pioneering neural architectures to shaping the global research community, the Deep Learning Group, led by Benjamin Erichson, PhD, has had numerous standout accomplishments.
With new research accepted at top-tier machine learning conferences, leadership in cross-institutional initiatives, and continued work at the intersection of scientific computing and AI, the group exemplifies ICSI’s mission of advancing foundational research to accelerate discovery.
Research Recognized at NeurIPS, ICML, ICLR, and AISTATS
It’s been a busy year for Dr. Erichson and his team. Five of the group’s major papers were recognized at leading machine learning conferences—NeurIPS, ICML, ICLR, and AISTATS—pushing forward our understanding of long-memory models, robust neural parameterizations, and security for Large Language Models (LLMs).
NeurIPS 2025: Improving Memory in Mamba with Block-Biased State Space Models
A new paper from ICSI’s Deep Learning Group, co-authored by Annan Yu and Benjamin Erichson, has been accepted to NeurIPS 2025, one of the premier conferences in machine learning. The paper, “Fixing Mamba’s Memory: Block-Biased State Space Models”, tackles a key challenge in long-sequence modeling using Mamba, a promising alternative to Transformers.
While Mamba has shown strong results, it struggles with benchmarks that require remembering information over long spans. Yu and Erichson diagnose the problem—fast memory decay, low model expressiveness, and unstable training—and introduce B2S6, a lightweight fix that uses block-level biases to stabilize and enhance memory retention. Their method improves performance on the Long-Range Arena benchmark and maintains strong results on language modeling, pushing the boundaries of what efficient, general-purpose sequence models can achieve.
ICML 2025: Strengthening Jailbreak Detection via Emojis
In a paper accepted to ICML 2025, Benjamin Erichson, alongside Zhipeng Wei and Yuqi Liu, introduced “Emoji Attack”—a study that exposes blind spots in LLM safety systems. The team found that inserting emojis into text can distort how “Judge LLMs” tokenize and interpret prompts, tricking them into misclassifying harmful content as safe.
By exploiting token segmentation bias, the attack demonstrates how even minor, human-readable changes can bypass AI safety filters. This work raises questions about the reliability of current moderation techniques and underscores the need for more robust defenses in the growing landscape of AI-generated content.
ICLR 2025: Spotlight on Frequency Bias and Long-Memory SSMs
Two Deep Learning Group papers were accepted to ICLR 2025, with “Tuning Frequency Bias of State Space Models” receiving a spotlight recognition. This work, led by Annan Yu, Dongwei Lyu, Soon Hoe Lim, Michael Mahoney, and Benjamin Erichson, tackles a common limitation in State Space Models (SSMs)—their tendency to focus too heavily on slow-changing (low-frequency) signals. The team proposed new training techniques, including initialization scaling and a frequency-aware gradient filter, that allow SSMs to better capture fast-changing, high-frequency patterns. These improvements significantly boost performance in tasks like image denoising and long-range sequence modeling.
The group also presented “HOPE for a Robust Parameterization of Long-Memory State Space Models,” introducing a new framework—HOPE—that leverages mathematical structures called Hankel operators to improve how memory is encoded and retained in SSMs. This approach leads to more stable training and better performance on long-sequence tasks, where retaining past information is critical. Together, these papers underscore ICSI’s leadership in pushing the boundaries of scientific deep learning—making AI systems more memory-efficient, mathematically grounded, and adaptable to complex real-world signals.
AISTATS 2025: Modeling Long-Term Dependencies with Delay-Driven RNNs
In their latest paper accepted to AISTATS 2025, “Gated Recurrent Neural Networks with Weighted Time-Delay Feedback,” Benjamin Erichson, along with co-authors Soon Hoe Lim and Michael Mahoney, developed a new kind of AI model that’s better at remembering information over time. Their method, called τ-GRU, adds a feedback system that lets the AI look back at what it learned earlier. This helps the model avoid a common problem in AI called “forgetting,” where important information gets lost in long sequences.
The team’s approach was inspired by how memory works in biology and shows strong results in tasks like image recognition and text understanding. Their model not only learns faster but also uses fewer resources—making it especially useful for scientific and real-world applications.
Co-Organizing Major Scientific Events
In addition to pursuing innovative research, Dr. Erichson is helping shape the future of AI in science through two major events.
Deep Learning for Science Summer School (DL4SCI)
Co-organized by Benjamin Erichson and Lawrence Berkeley National Laboratory, the DL4SCI Summer School is a week-long intensive program designed to bridge the gap between cutting-edge AI research and real-world scientific challenges. Held in June 2025, the school offered hands-on tutorials and deep dives into generative AI, LLMs, transformers, diffusion models, and more—all with a strong emphasis on practical application in high-performance computing environments. The event brought together students, postdocs, and researchers from around the world to collaborate, experiment, and build scalable AI solutions for science.
Berkeley Lab AI for Science Summit (BLASS 2024)
BLASS was a high-impact gathering of scientists, engineers, and AI experts exploring how artificial intelligence can accelerate breakthroughs in fields ranging from materials science to climate modeling. Co-organized by Benjamin Erichson and hosted at Lawrence Berkeley National Laboratory and venues across downtown Berkeley, California, the summit focused on cross-disciplinary collaboration and responsible innovation in AI-driven scientific computing. Sessions included research talks, panels on ethics and safety, and networking opportunities that fostered long-term partnerships across academia, industry, and national labs.
Looking Ahead: Scaling AI for Science
As scientific challenges grow more complex and AI tools become more powerful, the work of the Deep Learning Group is more relevant than ever. With foundational research in state space models, generative AI, and robust neural architectures, the group is positioned to leverage ICSI’s collaborative ecosystem to spearhead the next wave of innovation—building machine learning systems that are not just more powerful, but more interpretable, scalable, and aligned with scientific inquiry.
This story was published in January 2026 as part of a retrospective series highlighting ICSI’s accomplishments and impacts over the years. To learn about our ongoing work, explore our Core Research Themes.
