Tim Tsz-Kit Lau
|
AI Researcher Email: timlautk [AT] gmail.com |
About Me
I am currently an AI Researcher at DRW, based in Palo Alto, California.
I was a postdoctoral researcher at the University of Pennsylvania, jointly supervised by Prof. Weijie Su in Department of Statistics and Data Science, The Wharton School and, by courtesy, Departments of Computer and Information Science (CIS), Biostatistics, Epidemiology and Informatics, and Mathematics, and Prof. Qi Long in Department of Biostatistics, Epidemiology and Informatics, Perelman School of Medicine and, by courtesy, Departments of Statistics and Data Science and CIS, from November 2024 to May 2025.
Before that, I was a postdoctoral principal researcher in Econometrics and Statistics at the University of Chicago Booth School of Business, supervised by Prof. Mladen Kolar, from October 2023 to October 2024.
I received my Ph.D. in Statistics and Data Science from Northwestern University in 2023,
advised by Prof. Han Liu.
In my published research, I have worked broadly at the interface of statistics, machine learning and optimization, with applications to large language models (LLMs) and more broadly generative AI (GenAI). My current research aims to advance pre- (and post-)training science and scaling of large deep learning models through algorithmic and engineering perspectives of large-scale distributed stochastic nonconvex optimization methods, including architecture–optimizer co-design as well as developing efficient optimizers and batch-size strategies, in a theoretically principled way, in order to improve both pre-training and post-training scaling and training stability. See also my bio for details.
News
July 2026: We posted a reorganized and more formalized version of the below symmetry-compatible optimizers work for a mathematical optimization audience on Optimization Online.
May 2026: The paper Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers is out on arXiv. This paper introduces a symmetry-compatible principle for LLM optimizer design, leading to an end-to-end layerwise optimizer stack where every major matrix-valued parameter—including embeddings, LM heads, SwiGLU MLPs, and MoE routers—has its own principled update.
February 2026: Huge congratulations to my postdoc supervisor Prof. Weijie Su for winning the 2026 COPSS Presidents’ Award!
January 2026: The slides for the PolarGrad paper is available here.
May 2025: The paper PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective is out on arXiv.
March 2025: The paper Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism is accepted to the 2nd Conference on Parsimony and Learning (CPAL) Proceedings Track.
Other recent preprints on adaptive batch size schemes for large-scale distributed training: arXiv:2406.139366, arXiv:2402.11215