Song Han is a research director at NVIDIA, where he leads the Efficient AI team and co-leads the NVIDIA Singapore Lab. His research focuses on efficient AI computing. He received his PhD from Stanford University, advised by Prof. Bill Dally. Song proposed the "Deep Compression" technique, now widely used in efficient AI, and the "Efficient Inference Engine," which first brought weight sparsity to modern AI accelerator design. His team recently worked on LLM quantization, sparse attention, long-context inference, efficient visual generation and world models (research blogs). He received Best Paper Awards at ICLR'16, FPGA'17, and MLSys'24/26. Song was named to MIT Technology Review's "35 Innovators Under 35" for Deep Compression, which "lets powerful artificial intelligence (AI) programs run more efficiently on low-power mobile devices." He is also a recipient of the NSF CAREER Award for "efficient algorithms and hardware for accelerated machine learning," the IEEE "AI's 10 to Watch: The Future of AI" award, and the Sloan Research Fellowship.