Efficient AI Computing,
Transforming the Future.

Projects

To choose projects, simply check the boxes of the categories, topics and techniques.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference

IROS 2026
 (
)

VLASH is a general asynchronous inference framework for VLAs that delivers smooth, accurate, and low-latency control with no overhead or architectural changes. By rolling the robot state forward with the previous action chunk, it achieves up to 2.03× speedup and 17.4× lower reaction latency while fully preserving accuracy.

DeltaQuant: 4-bit Video Diffusion Models with Spatiotemporal Delta Smoothing

CVPR 2026
 (
)

A 4-bit weight-activation quantization method for video diffusion models that exploits spatiotemporal activation similarity, compressing Wan 2.2 by 2.9× and cutting memory by 2.3×.

ForeAct: Steering Your VLA with Efficient Visual Foresight Planning

CVPR 2026 Highlight
 (
Highlight
)

ForeAct is a plug-and-play visual foresight planner that enables state-of-the-art VLAs to anticipate high-fidelity future observations for improved decision-making, generating 640×480 predictions in just 0.33s on a single H100 GPU without any architectural changes.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation

arxiv 2026
 (
)

An offline post-training method for Large Reasoning Models that eliminates the need for live teacher during on-policy distillation, achieving a 4.0x higher training speedup while maintaining state-of-the-art performance.