Efficient AI Computing,
Transforming the Future.

Projects

To choose projects, simply check the boxes of the categories, topics and techniques.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization

ICML 2026
 (
)

A training-free KV-cache quantization framework for auto-regressive video diffusion, cutting KV memory by up to 7.0× with under 4% end-to-end latency overhead.

FourTune: Towards Fully 4-Bit Efficient Post-Training for Diffusion Models

ICML 2026
 (
)

An end-to-end W4A4G4 post-training framework for diffusion models, cutting memory 2.25× and boosting training throughput 2.27× over BF16 LoRA on FLUX.1-dev.

VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference

IROS 2026
 (
)

VLASH is a general asynchronous inference framework for VLAs that delivers smooth, accurate, and low-latency control with no overhead or architectural changes. By rolling the robot state forward with the previous action chunk, it achieves up to 2.03× speedup and 17.4× lower reaction latency while fully preserving accuracy.

DeltaQuant: 4-bit Video Diffusion Models with Spatiotemporal Delta Smoothing

CVPR 2026
 (
)

A 4-bit weight-activation quantization method for video diffusion models that exploits spatiotemporal activation similarity, compressing Wan 2.2 by 2.9× and cutting memory by 2.3×.