Efficient AI Computing,
Transforming the Future.

Bet Small to Win Big: Efficient 2K Video Generation via Deeper-compression AutoEncoder, Linear Attention and Two-Stage Refiner

TL;DR

This blog details SANA-Video transition to a Two-Stage Inference Paradigm. By leveraging a high-compression base model for structural discovery (Stage 1) and a step-distilled refiner for high-frequency detail injection (Stage 2), we achieve 2K resolution output without compromising the latency profile of the original 720p generation.

Related Projects

Linear Attention