Efficient AI Computing,
Transforming the Future.

TinyML and Efficient AI Computing

6.5940

Fall

2026

https://efficientml.ai

This course focuses on efficient machine learning and systems. This is a crucial area as deep neural networks demand extraordinary levels of computation, hindering its deployment on everyday devices and burdening the cloud infrastructure. This course introduces efficient AI computing techniques that enable powerful deep learning applications on resource-constrained devices. Topics include model compression, pruning, quantization, neural architecture search, distributed training, model serving, model parallelism, gradient compression, and on-device fine-tuning. It also introduces application-specific acceleration techniques for large language models and diffusion models. Students will get hands-on experience implementing model compression techniques and deploying large language models on a laptop.

  • Prerequisites: 6.191 Computation Structures and 6.390 Intro to Machine Learning. Students who don’t full-fill the prerequisites will be de-registered in the second week of class. If you believe you have equivalent prior experience (e.g., a computer architecture course taken during your undergraduate studies at another institution), you may petition for consideration.

Instructor

Associate Professor

Announcements

Currently no active announcements.

Schedule

Date

Lecture

Logistics

Introduction

Sep 10

Lecture
1
:

Introduction

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Basics of Deep Learning

Sep 15

Lecture
2
:

Basics of Deep Learning

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Lab 0 out

Chapter I: Efficient Inference

Sep 16

Lecture
3
:

Chapter I: Efficient Inference

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Pruning and Sparsity (Part I)

Sep 17

Lecture
3
:

Pruning and Sparsity (Part I)

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Pruning and Sparsity (Part II)

Sep 22

Lecture
4
:

Pruning and Sparsity (Part II)

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Lab 1 out

Quantization (Part I)

Sep 24

Lecture
5
:

Quantization (Part I)

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Lab 0 due

Quantization (Part II)

Sep 29

Lecture
6
:

Quantization (Part II)

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Neural Architecture Search (Part I)

Oct 1

Lecture
7
:

Neural Architecture Search (Part I)

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Lab 1 due, Lab 2 out

Neural Architecture Search (Part II)

Oct 6

Lecture
8
:

Neural Architecture Search (Part II)

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Knowledge Distillation

Oct 8

Lecture
9
:

Knowledge Distillation

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Student Holiday

Oct 13

Lecture
9
:

Student Holiday

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

MCUNet: TinyML on Microcontrollers

Oct 15

Lecture
10
:

MCUNet: TinyML on Microcontrollers

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Lab 2 due, Lab 3 out

TinyEngine and Parallel Processing

Oct 20

Lecture
11
:

TinyEngine and Parallel Processing

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Chapter II: Domain-Specific Optimization

Oct 21

Lecture
12
:

Chapter II: Domain-Specific Optimization

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Transformer and LLM

Oct 22

Lecture
12
:

Transformer and LLM

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

LLM Quantization and Deployment

Oct 27

Lecture
13
:

LLM Quantization and Deployment

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Lab 3 due, Lab 4 out

LLM Post Training

Oct 29

Lecture
14
:

LLM Post Training

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Project ideas out

Long Context LLM

Nov 3

Lecture
15
:

Long Context LLM

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Vision Transformer

Nov 5

Lecture
16
:

Vision Transformer

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Lab 4 due, Lab 5 out

GAN, Video, and Point Cloud

Nov 10

Lecture
17
:

GAN, Video, and Point Cloud

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Diffusion Model

Nov 12

Lecture
18
:

Diffusion Model

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Chapter III: Efficient Training

Nov 16

Lecture
19
:

Chapter III: Efficient Training

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Distributed Training (Part I)

Nov 17

Lecture
19
:

Distributed Training (Part I)

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Lab 5 due

Distributed Training (Part II)

Nov 19

Lecture
20
:

Distributed Training (Part II)

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Project proposal due

On-Device Training and Transfer Learning

Nov 24

Lecture
21
:

On-Device Training and Transfer Learning

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Thanksgiving

Nov 26

Lecture
21
:

Thanksgiving

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Chapter IV: Advanced Topics

Nov 30

Lecture
22
:

Chapter IV: Advanced Topics

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Advanced LLM and Agents

Dec 1

Lecture
22
:

Advanced LLM and Agents

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Final Project Presentation

Dec 3

Lecture
23
:

Final Project Presentation

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Final Project Presentation

Dec 8

Lecture
24
:

Final Project Presentation

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Final Project Presentation

Dec 10

Lecture
25
:

Final Project Presentation

[Slides]
[Slides]
[Video]
[Video]
[Video (Live)]
[Video (Live)]

Logistics

The class requirements include five labs, and one final project. This is a PhD level course, and by the end of this class you should have a good understanding of efficient deep learning techniques, and be able to deploy large language models (LLMs) on your laptop.

Note that this class does not have any tests or exams.

Labs

There will be 5 labs over the course of the semester.

  • Lab1: Pruning
  • Lab2: Quantization
  • Lab3: Neural architecture search
  • Lab4: LLM compression
  • Lab5: LLM deployment on laptop

Collaboration Policy

Labs must be done individually: each student must hand in their own answers. However, it is acceptable to collaborate when figuring out answers and to help each other solve the problems. We will be assuming that, as participants in a graduate course, you will be taking the responsibility to make sure you personally understand the solution arising from such collaboration. You also must indicate on each homework with whom you have collaborated.

Late Policy

You will be allowed 6 total homework late days without penalty for the entire semester. You may be late by up to 6 days on any homework assignment. Once those days are used, you will be penalized according to the following policy:

  • Homework is worth full credit at the due time on the due date.
  • The allowed late days are counted by day (i.e., each new late day starts at 11:59 pm ET).
  • Once the allowed late days are exceeded, the penalty is 50% per late day counted by day.
  • The homework is worth zero credit 2 days after exceeding the late day limit.

You must turn in at least 4 of the 5 assignments, even if for zero credit, in order to pass the course.

Final Project

The class project will be carried out in groups, and has three main parts:

  • proposal: choose from a list of suggested projects, or propose your own project
  • poster presentation
  • final report (4 pages, using the NeurIPS template)