Mathematical Machine Learning Theory
Rigorous properties of learning models, including optimisation, generalisation, asymptotic limits, and scaling behaviour.
University of Warwick · Mathematics
Undergraduate Student in Mathematics
My research focuses on the mathematical theory of machine learning, especially rigorous structures underlying neural-network training. Under the supervision of Fanghui Liu, I currently study the transfer of Warmup–Stable–Decay (WSD) learning-rate schedules under μP across network widths and training steps.
Mathematics · Machine Learning Theory
I am a Mathematics undergraduate at the University of Warwick, with a long-term research focus on machine learning theory. I am primarily interested in rigorous mathematical properties of learning systems, including neural-network theory, optimisation dynamics, infinite-width limits, hyperparameter transfer, and scaling laws.
Fanghui Liu is my research supervisor. Our current research investigates the transfer of three-parameter WSD learning-rate schedules under μP. A central idea is cumulative-learning-rate normalisation, which puts different step counts into a shared dynamical coordinate system and enables prediction of longer-run objectives from shorter reference runs, alongside transfer across model widths.
I am especially interested in problems that can be reduced, through probability, linear algebra, optimisation, and asymptotic analysis, to explicit mathematical structures that support rigorous and verifiable explanations of neural-network training.
Rigorous properties of learning models, including optimisation, generalisation, asymptotic limits, and scaling behaviour.
Parametrisation, feature learning, infinite-width limits, and the cross-width stability of hyperparameters.
WSD schedules and the transfer of their optimal parameters across network widths and training steps.
Optimisation, probability, spectral analysis, and asymptotic tools for training dynamics, stability, and deterministic limits.
The central idea of this project is a cumulative-learning-rate normalisation of WSD schedules. It places different training-step counts in a common dynamical coordinate system, while the final-to-peak ratio and decay fraction retain the schedule shape. This avoids directly comparing raw learning-rate parameters that no longer represent equivalent training across different step counts.
In these coordinates, we derive a finite-step expansion of the complete three-parameter WSD objective with a uniform second-order remainder. We rigorously show that one complete objective at a shorter step count is generally insufficient to determine the objective at a longer target, whereas two shorter-step reference objectives can predict the full target objective with second-order accuracy under suitable regularity conditions. We also connect cross-step prediction to fixed-step width stability under μP.
Rather than extrapolating only a scalar optimal peak learning rate, this project studies the complete objective over three-parameter WSD schedules: why one short-run reference is generally insufficient, and how two references can predict schedule performance at longer step counts.
The framework provides a mathematical explanation of why shorter-run information can predict the full WSD landscape at longer training budgets, offering a theoretical basis for reducing repeated target-scale tuning.
This project studies the global geometry of the wide-limit loss as a function of the learning rate in a two-step deep linear µP network. The central question is whether width-stable loss curves identify a unique optimal learning rate.
The results show that uniqueness is a stable generic phenomenon rather than an unconditional law. In exceptional cases, the correct limiting object is the full optimiser set.
I am a 4×4×4 speedcuber and have competed in official WCA competitions. My official results include a 34.22-second single and a 41.46-second average. My cubing profile is available here.
A 4×4×4 cube solve
Email: eric_feng2006@outlook.com