TinyML and Efficient AI Computing
This course focuses on efficient machine learning and systems. This is a crucial area as deep neural networks demand extraordinary levels of computation, hindering its deployment on everyday devices and burdening the cloud infrastructure. This course introduces efficient AI computing techniques that enable powerful deep learning applications on resource-constrained devices. Topics include model compression, pruning, quantization, neural architecture search, distributed training, model serving, model parallelism, gradient compression, and on-device fine-tuning. It also introduces application-specific acceleration techniques for large language models and diffusion models. Students will get hands-on experience implementing model compression techniques and deploying large language models on a laptop.
- Lecture Videos:https://live.efficientml.ai/
- Time:
Tuesday/Thursday 4:00-5:30 PM
- Location:54-100
- Office Hour:
Thursday 5:00-6:00 pm Eastern Time, 38-344 Meeting Room
- Discussion:Piazza (Coming Soon)
- Homework Submission:Canvas (Coming Soon)
- Contact:
- For external inquiries, personal matters, or emergencies, you can email us at efficientml-staff [at] mit.edu.
- Prerequisites: 6.191 Computation Structures and 6.390 Intro to Machine Learning. Students who don’t full-fill the prerequisites will be de-registered in the second week of class. If you believe you have equivalent prior experience (e.g., a computer architecture course taken during your undergraduate studies at another institution), you may petition for consideration.
Teaching Assistants
Announcements
Schedule
Date
Lecture
Logistics
Introduction
Sep 10
Introduction
Basics of Deep Learning
Sep 15
Basics of Deep Learning
Lab 0 out
Chapter I: Efficient Inference
Sep 16
Chapter I: Efficient Inference
Pruning and Sparsity (Part I)
Sep 17
Pruning and Sparsity (Part I)
Pruning and Sparsity (Part II)
Sep 22
Pruning and Sparsity (Part II)
Lab 1 out
Quantization (Part I)
Sep 24
Quantization (Part I)
Lab 0 due
Quantization (Part II)
Sep 29
Quantization (Part II)
Neural Architecture Search (Part I)
Oct 1
Neural Architecture Search (Part I)
Lab 1 due, Lab 2 out
Neural Architecture Search (Part II)
Oct 6
Neural Architecture Search (Part II)
Knowledge Distillation
Oct 8
Knowledge Distillation
Student Holiday
Oct 13
Student Holiday
MCUNet: TinyML on Microcontrollers
Oct 15
MCUNet: TinyML on Microcontrollers
Lab 2 due, Lab 3 out
TinyEngine and Parallel Processing
Oct 20
TinyEngine and Parallel Processing
Chapter II: Domain-Specific Optimization
Oct 21
Chapter II: Domain-Specific Optimization
Transformer and LLM
Oct 22
Transformer and LLM
LLM Quantization and Deployment
Oct 27
LLM Quantization and Deployment
Lab 3 due, Lab 4 out
LLM Post Training
Oct 29
LLM Post Training
Project ideas out
Long Context LLM
Nov 3
Long Context LLM
Vision Transformer
Nov 5
Vision Transformer
Lab 4 due, Lab 5 out
GAN, Video, and Point Cloud
Nov 10
GAN, Video, and Point Cloud
Diffusion Model
Nov 12
Diffusion Model
Chapter III: Efficient Training
Nov 16
Chapter III: Efficient Training
Distributed Training (Part I)
Nov 17
Distributed Training (Part I)
Lab 5 due
Distributed Training (Part II)
Nov 19
Distributed Training (Part II)
Project proposal due
On-Device Training and Transfer Learning
Nov 24
On-Device Training and Transfer Learning
Thanksgiving
Nov 26
Thanksgiving
Chapter IV: Advanced Topics
Nov 30
Chapter IV: Advanced Topics
Advanced LLM and Agents
Dec 1
Advanced LLM and Agents
Final Project Presentation
Dec 3
Final Project Presentation
Final Project Presentation
Dec 8
Final Project Presentation
Final Project Presentation
Dec 10
Final Project Presentation
Logistics
The class requirements include five labs, and one final project. This is a PhD level course, and by the end of this class you should have a good understanding of efficient deep learning techniques, and be able to deploy large language models (LLMs) on your laptop.
Note that this class does not have any tests or exams.
Labs
There will be 5 labs over the course of the semester.
- Lab1: Pruning
- Lab2: Quantization
- Lab3: Neural architecture search
- Lab4: LLM compression
- Lab5: LLM deployment on laptop
Collaboration Policy
Labs must be done individually: each student must hand in their own answers. However, it is acceptable to collaborate when figuring out answers and to help each other solve the problems. We will be assuming that, as participants in a graduate course, you will be taking the responsibility to make sure you personally understand the solution arising from such collaboration. You also must indicate on each homework with whom you have collaborated.
Late Policy
You will be allowed 6 total homework late days without penalty for the entire semester. You may be late by up to 6 days on any homework assignment. Once those days are used, you will be penalized according to the following policy:
- Homework is worth full credit at the due time on the due date.
- The allowed late days are counted by day (i.e., each new late day starts at 11:59 pm ET).
- Once the allowed late days are exceeded, the penalty is 50% per late day counted by day.
- The homework is worth zero credit 2 days after exceeding the late day limit.
You must turn in at least 4 of the 5 assignments, even if for zero credit, in order to pass the course.
Final Project
The class project will be carried out in groups, and has three main parts:
- proposal: choose from a list of suggested projects, or propose your own project
- poster presentation
- final report (4 pages, using the NeurIPS template)





