Learning better from your own predictions: Self-distillation as optimal spectral shrinkage
In modern machine learning pipelines, distillation has emerged as a promising approach to improving model performance without requiring additional training data. Yet a fundamental statistical question remains: When does self-distillation reduce out-of-sample risk or improve estimation relative to classical statistical methods? In this talk, I will introduce a rigorous framework for analyzing a broad class of high-dimensional estimators, which we call spectral-shrinkage estimators. Under spiked covariance models, we show that self-distillation achieves optimal prediction and estimation performance within this class, strictly outperforming ridge regression, minimum-norm interpolation, and principal components regression. More precisely, when the covariance has r distinct spikes, r-step self-distillation attains the optimal spectral-shrinkage rule, and, in general, r steps are necessary. This behavior contrasts sharply with the isotropic setting, where optimally tuned ridge regression is the best spectral-shrinkage estimator. Time permitting, I will also discuss a federated learning setting where multiple data centers transmit locally computed estimators to a central server for aggregation. We characterize the optimal local shrinkage rules and aggregation weights, and show that the best local rule can again be implemented through self-distillation—though it differs from the optimal rule in the centralized setting. Together, these results demonstrate when and why self-distillation improves statistical performance, and place it within a broader framework connecting modern distillation procedures with classical shrinkage methods.