Optimization ∩ AI · Theory ∩ Practice
My full name is Hao-Jun Michael Shi, although I normally go by Michael. I currently lead a small team developing and translating training-algorithm research into production systems for ranking and recommendation and LLM pre-training applications. This includes our distributed PyTorch implementation of Shampoo and other matrix optimizers, research into semi-synchronous training and GPA, and the design of metrics and techniques for batch size selection and training stability. My background is in numerical optimization, and I have previously designed algorithms for stochastic, noisy, and derivative-free optimization, contributed to Facebook's open-source deep learning recommendation model (DLRM), and developed embedding compression techniques.
Outside of work, I enjoy eating, reading, dabbling in board games and video games, playing basketball, rooting for the Golden State Warriors and Valkyries and UCLA Bruins, spending time with my family, and serving at my church.
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent accepted at ICML 2026 (Seoul). Thanks to Params and Hiroki for presenting on the team's behalf!
Presented Re-Thinking DiLoCo through Generalized Primal Averaging for Distributed Training, Digital Futures Seminar, KTH Royal Institute of Technology (Stockholm). Thanks to Mikael Johansson for the invitation!
Thesis opponent for Zesen Wang's licentiate seminar at KTH (Stockholm).
Presented Beyond AdamW: Developments in Training Algorithms for Deep Learning, Stanford MS&E. Thanks to Madeleine Udell for the invitation!
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner accepted and presented at NeurIPS 2025 (San Diego).
1st Place (External Tuning Ruleset), MLCommons AlgoPerf benchmark.
“So, whether you eat or drink, or whatever you do, do all to the glory of God.”— 1 Corinthians 10:31