Deep learning has revolutionized our field, leading to so many new and exceptionally successful applications of ML, especially with language and image modeling. Deep learning is also extremely resource intensive. The scaling up of deep architectures has precipitated the huge successes, but is it really necessary? We still don’t know. The resulting highly nonconvex optimization problems require careful training based on many heuristics that people have come to rely on — including the use of certain optimizers, step size choices, and overparameterization, all with the goal of improving the nonconvex landscape.
My group studies ways of solving the same optimization problems but without massive overparameterization. We are very interested in swapping out compute- or memory-intensive architecture blocks with approximations that use compact GPU-friendly representations, such as low-rank and block-sparse matrices. Our goal is to both provide novel approaches that don’t require massive resources at the same time as we build a deeper mathematical understanding of when efficient deep learning is possible.
- Balzano, L., Ding, T., Haeffele, B. D., Kwon, S. M., Qu, Q., Wang, P., Wang, Z., & Yaras, C. (2026). “An Overview of Low-Rank Structures in the Training and Adaptation of Large Models.” (arXiv:2503.19859). Accepted to Signal Processing Magazine, special issue on the Mathematics of Deep Learning.
- Yaras, C., Xu, A., Abillama, P., Lee, C., & Balzano, L. (2025). “MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention.” Advances in Neural Information Processing Systems, 38, 3445–3471.
- Yaras, C., Wang, P., Balzano, L., & Qu, Q. (2024, June 6). “Compressible Dynamics in Deep Overparameterized Low-Rank Learning & Adaptation.” Forty-first International Conference on Machine Learning, oral presentation.
- Kwon, S. M., Zhang, Z., Song, D., Balzano, L., & Qu, Q. (2024). “Efficient Low-Dimensional Compression of Overparameterized Models.” Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, 1009–1017.