Efficient AI Computing,
Transforming the Future.

Projects

To choose projects, simply check the boxes of the categories, topics and techniques.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

LongLive: Real-time Interactive Long Video Generation

ICLR 2026
 (
)

We present LongLive, a frame-level autoregressive (AR) framework for real-time and interactive long video generation. Try LongLive: turning your interactive prompts into long videos-instantly, as you type!

QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs

ICLR 2026
 (
)

QeRL is a framework that accelerates LLM training by combining NVFP4 quantization with LoRA. It reduces memory overhead, allows 32B model training on a single H100 GPU, and improves performance by leveraging quantization noise to

Fast-dLLM v2: Efficient Block-Diffusion Large Language Model

ICLR 2026
 (
)

Our approach introduces a novel training recipe that combines a block diffusion mechanism with a complementary attention mask, enabling blockwise bidirectional context modeling without sacrificing AR training objectives.

Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding

ICLR 2026
 (
)

We introduce a novel block-wise approximate KV Cache mechanism tailored for bidirectional diffusion models, enabling cache reuse with negligible performance drop.