Neural Architecture Search (NAS)

Projects

Tiny Machine Learning Projects

NeurIPS 2020/2021/2022, MICRO 2023, ICML 2023, MLSys 2024, IEEE CAS Magazine 2023

This TinyML project aims to enable efficient AI computing on the edge by innovating model compression techniques as well as high-performance system design.

Tiny Machine Learning: Progress and Futures [Feature]

We discuss the definition, challenges, and applications of TinyML.

On-Device Training Under 256KB Memory

NeurIPS 2022

(

)

In MCUNetV3, we enable on-device training under 256KB memory, using less than 1/1000 memory of PyTorch while matching the accuracy on the visual wake words application using system-algorithm co-design.

Lite Pose: Efficient Architecture Design for 2D Human Pose Estimation

CVPR 2022

(

)

Litepose is an efficient neural network architecture for 2D human pose estimation.

NAAS: Neural Accelerator Architecture Search

DAC 2021

(

)

As a data-driven approach, NAAS holistically composes highly matched accelerator and neural architectures together with efficient compiler mapping.

MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning

NeurIPS 2021

(

)

In MCUNetV2, we propose a generic patch-by-patch inference scheduling, which operates only on a small spatial region of the feature map and significantly cuts down the peak memory. We further propose network redistribution to shift the receptive field and FLOPs to the later stage and reduce the computation overhead.

Blog Posts

On-Device Training Under 256KB Memory

November 28, 2022

In MCUNetV3, we enable on-device training under 256KB SRAM and 1MB Flash, using less than 1/1000 memory of PyTorch while matching the accuracy on the visual wake words application. It enables the model to adapt to newly collected sensor data and users can enjoy customized services without uploading the data to the cloud thus protecting privacy.

Reducing the carbon footprint of AI using the Once-for-All network

July 3, 2020

“The aim is smaller, greener neural networks,” says Song Han, an assistant professor in the Department of Electrical Engineering and Computer Science. “Searching efficient neural network architectures has until now had a huge carbon footprint. But we reduced that footprint by orders of magnitude with these new methods.”

Auto Hardware-Aware Neural Network Specialization on ImageNet in Minutes

July 2, 2020

This tutorial introduces how to use the Once-for-All (OFA) Network to get specialized ImageNet models for the target hardware in minutes with only your laptop.

Efficiently Understanding Videos, Point Cloud and Natural Language on NVIDIA Jetson Xavier NX

May 22, 2020

Thanks to NVIDIA’s amazing deep learning eco-system, we are able to deploy three applications on Jetson Xavier NX soon after we receive the kit, including efficient video understanding with Temporal Shift Module (TSM, ICCV’19), efficient 3D deep learning with Point-Voxel CNN (PVCNN, NeurIPS’19), and efficient machine translation with hardware-aware transformer (HAT, ACL’20).