Hao-Jun Michael Shi

Research Scientist — FAIR, Meta Superintelligence Labs
Menlo Park, CA

Optimization ∩ AI · Theory ∩ Practice

Portrait of Hao-Jun Michael Shi

About

My full name is Hao-Jun Michael Shi, although I normally go by Michael. I currently lead a small team developing and translating training-algorithm research into production systems for ranking and recommendation and LLM pre-training applications. This includes our distributed PyTorch implementation of Shampoo and other matrix optimizers, research into semi-synchronous training and GPA, and the design of metrics and techniques for batch size selection and training stability. My background is in numerical optimization, and I have previously designed algorithms for stochastic, noisy, and derivative-free optimization, contributed to Facebook's open-source deep learning recommendation model (DLRM), and developed embedding compression techniques.

Outside of work, I enjoy eating, reading, dabbling in board games and video games, playing basketball, rooting for the Golden State Warriors and Valkyries and UCLA Bruins, spending time with my family, and serving at my church.

Bio: education, positions, and honors →

Research Interests

·  Numerical Optimization ·  Deep Learning ·  Mathematical Software ·  Scientific Computing

News & Highlights

Selected Publications

ICML 2026
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
H. Naganuma, S. Gupta, Y. Briki, I. Mitliagkas, I. Rish, P. Raman, H.-J.M. Shi
Preprint 2026
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
R. Eschenhagen, A. Cai, T.-H. Lee, H.-J.M. Shi
NeurIPS 2025
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
R. Eschenhagen, A. Defazio, T.-H. Lee, R. E. Turner, H.-J.M. Shi
Preprint 2025
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
A. Defazio, K. Mishchenko, P. Raman, H.-J.M. Shi, L. Xiao
Tech Report 2023
A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
H.-J.M. Shi, T.-H. Lee, S. Iwasaki, J. Gallego-Posada, Z. Li, K. Rangadurai, D. Mudigere, M. Rabbat
KDD 2020
Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation Systems
H.-J.M. Shi, D. Mudigere, M. Naumov, J. Yang
ICML 2018
A Progressive Batching L-BFGS Method for Machine Learning
R. Bollapragada, D. Mudigere, J. Nocedal, H.-J.M. Shi, P. T. P. Tang

Full publication list on Google Scholar →

Software

PyTorch Distributed Shampoo — A PyTorch implementation of the Distributed Shampoo optimizer compatible with DDP and FSDP.
PyTorch-LBFGS — A modular PyTorch implementation of stochastic L-BFGS.
“So, whether you eat or drink, or whatever you do, do all to the glory of God.”— 1 Corinthians 10:31