·
AI & ML interests
Training dynamics, optimization & algorithm–hardware co-design
Recent Activity
Organizations
view article DeepSeek Engram × OLMo-core: Distributed Implementation
bird-of-paradise
• • 3
view article The Three Horsemen of Numerical Divergence in Hybrid Models
bird-of-paradise
• • 3
view article FlashAttention, Streaming Algorithms, and Numerical Stability in Modern ML Systems
bird-of-paradise
• • 1
view article Hopper: The Optimizer That Learns Parallelism 2x Faster Than Adam
bird-of-paradise
• • 4
view article Field Notes: Why Muon "Hollows Out" in RL (and What We Plan To DO Next)
bird-of-paradise
• • 6
view article Field Notes: The Dilemma of Training Reasoning with Muon
bird-of-paradise
• • 2
view article Scaling Is Not Plug-and-Play: What Muon Teaches Us About Optimizers at Scale
bird-of-paradise
• • 2
view article Beyond the Wrapper: Building High-Throughput Reasoning Agents with Async Kernels
bird-of-paradise
• • 1
view article Reproducing and Validating Distributed Muon 🐢✨: A Practical Verification of Communication Efficiency Claims
bird-of-paradise
• • 2