AI practitioner, mountaineer, and coffee nerd
Sheng Zha 查晟
LLM pre-training and algorithm-system co-design
Built teams, algorithms, and systems behind Amazon Nova, SageMaker HyperPod, Apache MXNet, and GluonNLP.

What I’m thinking about
Recent writing
Compute-Optimal Is Not Cluster-Optimal
MOSAIC jointly selects a sparse-MoE architecture, token budget, and parallel layout under a fixed cluster and training window.
Research Problems in Pretraining
A practitioner's account of what pretraining research can predict, where current methods break, and which questions remain open.
Your Org Has the Same Scaling Problem as a Badly Tuned Training Run
AI raised individual throughput but coordination overhead stayed fixed. For many product-engineering orgs, the bottleneck flipped from compute-bound to communication-bound.