Loading Events

Hosted by Mark Peletier

Speaker

Tom Jacobs, CISPA (Helmholtzcenter) Saarbrücken

Title

Controlling Implicit Regularization in Deep Learning via Weight Decay and Mirror Descent

Abstract

Classical learning theory predicts that overparameterized models should overfit, yet deep neural networks generalize well in this regime. A possible explanation for this is implicit regularization: gradient-based optimization biases solutions toward low-complexity structures (e.g., sparsity or low rank) even without explicit constraints, as observed in settings such as matrix sensing and attention models. In this seminar, I show that weight decay controls this bias: beyond its explicit role as L2-regularization, it modifies the optimization geometry (mirror map), effectively shifting the implicit regularization toward L1-type behavior and thereby promoting sparsity. By turning off weight decay during training, only the implicit effect remains, leading to better generalization. Leveraging this perspective, I introduce PILoT (Parametric Implicit Lottery Ticket), a sparsification method that exploits overparameterization and the L2-to-L1 transition in implicit regularization to produce sparse networks with minimal performance degradation. Building on these insights, I further introduce HAM (Hyperbolic Aware Minimization), a lightweight optimization method that captures the sparsity-inducing implicit bias using mirror descent, thereby directly controlling the implicit bias and leading to improved standard training and state-of-the-art performance in finding sparse networks.

Go to Top