Researchers have estimated that the size of AI models grows 240-fold every two years, while the memory capacity of the hardware to support all that computation merely doubles in the same timeframe. As AI increasingly finds its way into our homes, cities, and daily lives, AI models must be compressed in order to be practical for fast-responding, energy-efficient Internet of Things devices. But compressing AI models is a computational challenge that itself takes time and energy.
This is the challenge that inspired graduate student Peter (Jiayun) Wang and PI Stella Yu to create a new end-to-end solution to model compression, in work funded by Berkeley AI Research and the Berkeley Deep Drive project. Rather than focusing on removing unused parameters from a large model, their approach aims to squeeze more information onto fewer parameters, making models more efficient by building them smaller in the first place. The method outperforms previous solutions in terms of time and resource efficiency, and can even handle large-scale AI models such as transformer networks. In one paper, the team demonstrated the ability to compress AI models by 10 times while retaining accuracy. In another paper, they demonstrated that their proposed efficient AI models are robust to adversarial attacks.
What made ICSI a good place to pursue these projects?
“ICSI provides a platform with multifaceted support and many fellows sharing the same interest. This has greatly benefitted my professional growth and career, and I would highly recommend ICSI to others.”
Peter Wang
Former Vision Group Member
This story was published in January 2026 as part of a retrospective series highlighting ICSI’s accomplishments and impacts over the years. To learn about our ongoing work, explore our Core Research Themes.
