Research
Research Interests
In my academic research, I have worked broadly on the mathematical and statistical foundations of machine learning and artificial intelligence, with a more recent additional emphasis on addressing real engineering challenges of scaling AI architectures and algorithms. My current research aims to advance pre- (and post-)training science and scaling of large deep learning models through algorithmic and engineering perspectives of large-scale distributed stochastic nonconvex optimization methods, including architecture–optimizer co-design as well as developing efficient optimizers and batch-size strategies, in a theoretically principled way, in order to improve both pre-training and post-training scaling and training stability.
In particular, I am interested in
Advancing pre- (and post-)training science and scaling of large deep learning models through algorithmic and engineering perspectives of large-scale distributed stochastic nonconvex optimization methods, including architecture–optimizer co-design, efficient optimizers and batch-size strategies
Theory and applications of optimization and sampling techniques to generative AI (GenAI), e.g., efficient (pre- and post-)training of LLMs and mixture-of-expert models
The interplay between optimization and sampling
High-dimensional structured matrix estimation problems and its applications
Funding and Grants
The University of Chicago Data Science Institute — AI + Science Research Initiative:
Project Support Funds (Principal Investigator) with $20,000 equivalent of GPU compute (2024)
Project Title: Advancing state-of-the-art large-scale distributed training methods in the era of generative AI
Academic Services
Reviewer for
Journals
IEEE Transactions on Signal Processing (2018—1, 2019—1, 2020—2, 2025—1, 2026-1)
Neural Computation (2021)