Machine Learning Theory
Rigorous mathematical properties of learning models, including provability, generalisation, optimisation, and scaling laws.
University of Warwick · Mathematics
Undergraduate Student in Mathematics
My research direction is machine learning theory, with an emphasis on rigorous mathematical structures underlying learning systems. I have completed a theoretical research project on µP learning-rate transfer under the supervision of Fanghui Liu and am now preparing for the next stage of advanced research.
Mathematics · Machine Learning Theory
I am a Mathematics undergraduate at the University of Warwick, with a long-term research focus on machine learning theory. I am primarily interested in rigorous mathematical properties of learning systems, including neural-network theory, optimisation dynamics, asymptotic behaviour, and scaling laws.
Fanghui Liu is now my formal research supervisor. We have completed a theoretical research project on learning-rate transfer under µP and are preparing the next stage of advanced research. The specific topic will be added once it is determined.
I am especially interested in problems that can be reduced, through probability, linear algebra, optimisation, and asymptotic analysis, to explicit mathematical structures that support rigorous and verifiable explanations of neural-network training.
Rigorous mathematical properties of learning models, including provability, generalisation, optimisation, and scaling laws.
Parametrisation, representation learning, infinite-width limits, and the mathematical structure of training dynamics.
Mathematical relations among optimisation algorithms, hyperparameters, training trajectories, and stability.
Probability, spectral analysis, and asymptotic tools for deterministic limits and fluctuations in high-dimensional learning systems.
The next stage of advanced research is being prepared. The topic and research details will be added once they are determined.
This project studies the global geometry of the wide-limit loss as a function of the learning rate in a two-step deep linear µP network. The central question is whether width-stable loss curves identify a unique optimal learning rate.
The results show that uniqueness is a stable generic phenomenon rather than an unconditional law. In exceptional cases, the correct limiting object is the full optimiser set.
A study of generic uniqueness, structured nonuniqueness, and perturbative stability of optimal learning rates in a two-step deep linear µP network.
View paper PDF →I am a 4×4×4 speedcuber and have competed in official WCA competitions. My official results include a 34.22-second single and a 41.46-second average. My cubing profile is available here.
A 4×4×4 cube solve
Email: eric_feng2006@outlook.com