Efficient AI Computing,
Transforming the Future.

Projects

To choose projects, simply check the boxes of the categories, topics and techniques.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation

ICLR 2026 Oral
 (
Oral
)

We introduce Locality-aware Parallel Decoding to accelerate autoregressive image generation and achieve 13× faster than traditional AR models and at least 3.4× faster than previous parallelized AR models.

StreamingVLM: Real-Time Understanding for Infinite Video Streams

ICLR 2026
 (
)

StreamingVLM enables real-time understanding of infinite videos with low, stable latency. By aligning training on overlapped video chunks with an efficient KV cache, it runs at 8 FPS on a single H100. It achieves a 66.18% win rate vs. GPT-4o mini on a new benchmark with videos averaging over 2 hours long.

LongLive: Real-time Interactive Long Video Generation

ICLR 2026
 (
)

We present LongLive, a frame-level autoregressive (AR) framework for real-time and interactive long video generation. Try LongLive: turning your interactive prompts into long videos-instantly, as you type!

QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs

ICLR 2026
 (
)

QeRL is a framework that accelerates LLM training by combining NVFP4 quantization with LoRA. It reduces memory overhead, allows 32B model training on a single H100 GPU, and improves performance by leveraging quantization noise to