Efficient AI Computing,
Transforming the Future.

Team

Principal Investigator

Song Han is an associate professor at MIT EECS. He earned his PhD from Stanford, pioneering efficient AI computing techniques such as “Deep Compression” (pruning, quantization) and the “Efficient Inference Engine,” which first introduced weight sparsity to modern AI chips, making it one of the top-5 most cited papers in the 50-year history of ISCA (1953-2023). His innovations, including TinyML and hardware-aware neural architecture search (Once-for-All Network), have advanced AI model deployment on resource-constrained devices. His recent work on LLM quantization/acceleration (SmoothQuant, AWQ, StreamingLLM) has improved efficiency in LLM inference, adopted by NVIDIA TensorRT-LLM. Song received best paper awards at ICLR'16, FPGA'17, and MLSys'24, the NSF CAREER Award, “35 Innovators Under 35,” IEEE “AI’s 10 to Watch,” and the Sloan Research Fellowship. He developed the open lecture series EfficientML.ai to share advances in efficient ML research.

Team Members

Postdoctoral

His research focuses on building efficient machine learning systems, particularly in the areas of training, serving, and scheduling for foundation models. His work received the Distinguished Paper Award at ASPLOS. He is also a recipient of the Google PhD Fellowship and has been recognized as one of the ML and Systems Rising Stars by MLCommons.

Ph.D

Jiaming Tang is a third-year Ph.D. student at MIT, advised by Prof. Song Han. He was a member of ACM Honors Class, Shanghai Jiao Tong University. His research interests lie in efficient systems and algorithms for large language models. His work AWQ receives the Best Paper Award at MLSys 2024 and has been integrated into Transformers, SGLang, vLLM, FastChat, TensorRT-LLM, and TGI.

Junxian Guo is a second-year Ph.D. student at MIT, advised by Prof. Song Han. He was a member of ACM Honors Class, Shanghai Jiao Tong University. His research interests lie in systems and algorithms for long context models.

Shang Yang is a fourth-year Ph.D. student at MIT, advised by Prof. Song Han. He received his B.Eng. degree from Tsinghua University. His research focuses on efficient machine learning systems. He has led and co-led projects including QServe (MLSys'25), LServe (MLSys'25), AWQ (MLSys'24 Best Paper), TorchSparse++ (MICRO'23). His work received over 5,000 GitHub stars and more than 2,500 citations on Google Scholar.

Wenkun He is a first-year Ph.D. student at MIT, supervised by Prof. Song Han. He received his Bachelor's degree from Yao Class, Tsinghua University. His research focuses on efficient physical AI (World Action Models, Vision-Language-Action Models, etc.). He is also interested, and has relevant experience in efficient visual generation.

Xingyang Li is a first-year Ph.D. student at MIT, advised by Prof. Song Han. He received his B.Eng. degree from Shanghai Jiao Tong University, where he was a member of the ACM Honors Class. His research focuses on efficient machine learning systems, particularly accelerating generative AI with low-bit quantization and sparsity.

Yecheng Wu is a first-year Ph.D. student at MIT, supervised by Prof. Song Han. He received his Bachelor's degree from Yao Class, Tsinghua University. His research focuses on post-training of large language models. He also has experience in visual generative models and unified foundation models.

Zhuoyang Zhang is a third-year Ph.D. student at MIT, advised by Prof. Song Han. He received his Bachelor’s degree from Yao class, Tsinghua University. His research focuses on efficient machine learning algorithms and systems. His research has been recognized with oral and highlight presentations at ICLR and CVPR. His work has received over 4,000 stars on GitHub and over 2,500 citations on Google Scholar.

Master

Currently None.

Undergraduate

Currently None.

Graduated

Ph.D

His research focuses on the development of high-performance and efficient hardware architectures and software systems for deep learning. Zhekai leads the low-level architecture design of multiple hardware projects, including SpArch (HPCA'20), SpAtten (HPCA'21), PointAcc (MICRO'21), and LEGO (HPCA'25), which have received over 700 citations. Zhekai also leads the system and CUDA kernel development of Nunchaku, an efficient inference engine used by software projects including SVDQuant and SANA.

Ph.D

His research interest is in the intersection of machine learning, system, and computer graphics. He is currently working on building efficient and hardware-friendly generative models with its applications in computer vision and graphics. His work GAN Compression receives 1.1K stars on GitHub.

Ph.D

His research interests focus on the development of efficient algorithms and systems for deep learning, specifically large foundation models. His work has received over 8000 stars on GitHub. His work has a real-world impact: SmoothQuant has been integrated into NVIDIA's TensorRT-LLM, FasterTransformer and Intel's NeuralCompressor and is utilized in the LLMs of industry companies like Amazon, Meta, and Huggingface. StreamingLLM has been integrated into NVIDIA's TensorRT-LLM, Huggingface's transformers, and Intels' Extension for Transformers.

Ph.D

Yujun Lin graduated from MIT HAN Lab in April 2025. He joined NVIDIA Research as a research scientist after graduation. His research focuses on the intersection of computer architecture and machine learning, particularly the co-design of software and hardware for deep learning and its applications. Yujun was awarded the 2021 Qualcomm Innovation Fellowship, and he is the founding member of the new course on TinyML and efficient deep learning computing (MIT 6.S965) teaching crew, which received 12k views on YouTube.

Ph.D

Haotian is currently a research scientist at Meta. He received his PhD at MIT EECS with Prof. Song Han in January 2025. His research interests lie at the intersection of computer systems and machine learning. He is currently working on efficient multi-modal generation with foundation models. He has authored multiple papers with over 3,800 citations. Haotian has successfully advised several undergraduate students, and the intern students he mentored have continued as PhD students at MIT and UC Berkeley.

Postdoctoral

His research focuses on efficient deep learning, TinyML, embedded systems, and memory/storage systems. Wei-Chen has received several accolades for his work, including the MLSys Best Paper Award, the Best Poster Award at the NSF Athena AI Institute, the ACM/IEEE CODES+ISSS Best Paper Award, and the IEEE NVMSA Best Paper Award. In addition, he received first place (among 150 teams) in the flash consumption track of the ACM/IEEE TinyML Design Contest at ICCAD 2022. His research has received over 4,000 stars on GitHub, and his work "On-device training under 256KB memory" (MCUNetV3) was highlighted by the MIT homepage. He is the cofounder Eigen AI.

Ph.D

Hanrui Wang graduated from MIT HAN Lab in 2024 and cofounded Eigen AI. His research focuses on efficient AI and emerging hardware (e.g. quantum architecture). His research has been recognized by ACM student research competition 1st place award, best poster award at NSF AI Institute, Best Presentation Award as a DAC Young Fellow and appears in top conferences such as NeurIPS, ISCA, MICRO, HPCA, and DAC. His co-authored papers received ICML RL4RL Best Paper Award and QCE Best Paper Award. He is the recipient of Qualcomm Fellowship, Unitary Fund, and Nvidia Fellowship Finalist. He is the creator of TorchQuantum library which has been adopted by IBM and PyTorch Ecosystems. He is also the co-founder of QuCS lecture series for quantum education. Hanrui received his B. Eng. degree from Fudan University.

Ph.D

Han Cai graduated from MIT HAN Lab in May 2024. He joined NVIDIA Research as a research scientist after graduation. His research focuses on algorithms and acceleration of efficient deep learning computing. Han has made significant contributions to the field, including his work on hardware-aware neural architecture search (ProxylessNAS, Once-for-All), which has been integrated into PytorchHub@Meta, AutoGluon@Amazon, NNI@Microsoft, SONY Neural Architecture Search Library, SONY Model Compression Toolkit, and ADI Model Training and Synthesis Tool. His research has received 6.9K+ citations on Google Scholar and 5.2K+ stars on GitHub.

Ph.D

Zhijian graduated from MIT HAN Lab in May 2024. Zhijian joined UCSD as a tenure-track assistant professor, after a gap year at NVIDIA Research. His research focuses on efficient machine learning and systems. His work has been featured in oral and spotlight presentations at conferences such as NeurIPS, ICLR, and CVPR. He received the Qualcomm Innovation Fellowship. He was recognized as a Rising Star in ML and Systems by MLCommons and a Rising Star in Data Science by UChicago and UCSD. His work has received over 9000 citations on Google Scholar and over 11000 stars on GitHub. He received his B.Eng. degree from Shanghai Jiao Tong University.

Ph.D

Ji Lin graduated from MIT HAN Lab in Dec. 2023 and joined OpenAI as a research scientist. His research focuses on efficient deep learning computing, systems for ML and recently, accelerating large language models (LLMs). Ji is pioneering the research in the field of TinyML. His research has received over 10,000 citations on Google Scholar and over 8,000 stars on GitHub. His work on LLM quantization (AWQ) received the best paper award at MLSys'24. AWQ has been widely adopted by NVIDIA, Intel, Microsoft, AMD, HuggingFace, Berkeley to accelerate LLM inference. AWQ-quantized LLMs have been downloaded by more than 6 million times on HuggingFace. Ji is an NVIDIA Graduate Fellowship Finalist in 2020, and Qualcomm Innovation Fellowship recipient in 2022. His work has been covered by MIT Tech Review, MIT News (twice on MIT homepage and four times on MIT News), WIRED, Engadget, VentureBeat, etc.

Postdoctoral

Wei-Ming Chen is a Postdoctoral Associate at MIT EECS advised by Professor Song Han. His research focuses on TinyML, embedded systems, and real-time systems, with a particular emphasis on enabling efficient deep learning on Internet of Things (IoT) devices, such as microcontrollers. Chen's recent work on the MCUNet series (MCUNet, MCUNetv2, and MCUNetv3) has enabled efficient inference and training on devices with limited memory through the co-design of systems and algorithms. He is also a key contributor and maintainer of TinyEngine, an open-source library for high-performance and memory-efficient deep learning on microcontrollers. His work "On-device training under 256KB memory" (MCUNetV3) is highlighted by the MIT homepage in fall 2022. He received first place (among 150 teams) in the flash consumption track of the ACM/IEEE TinyML Design Contest at ICCAD 2022. He developed TinyChatEngine that enables LLM inference on the edge (laptop, Paspberry PI). His research has received more than 1,000 stars on GitHub. After graduation, he joined NVIDIA as a senior deep learning engineer working on large language model acceleration.

Master

Kevin Shao was an M.Eng student at MIT HAN Lab, working on autonomous driving and efficient 3D deep learning. After graduation, he joined Two Sigma.

Master

Driss Hafdi was an M.Eng student at MIT EECS, working on specialized hardware for mixed-precision quantization. After graduation, he joined Hudson River Trading.

Openings

Our lab is currently full in capacity.
Please do no email us regarding admissions.

Sponsors

We actively collaborate with industry partners on efficient AI, model compression and acceleration, LLM inference, physical AI.