Efficient AI Computing,
Transforming the Future.

Projects

To choose projects, simply check the boxes of the categories, topics and techniques.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition

Arxiv 2026
 (
)

Lightning Weave improves the accuracy-efficiency frontier of large reasoning models by composing capabilities from accuracy- and efficiency-oriented teachers into a single student through offline distillation.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models

Arxiv 2026
 (
)

Lightning OPD 2.0 enables effective cross-teacher on-policy distillation by mitigating style bias through cross-fitted residualization, allowing the SFT data generator and distillation teacher to be chosen independently.

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization

ICML 2026
 (
)

A training-free KV-cache quantization framework for auto-regressive video diffusion, cutting KV memory by up to 7.0× with under 4% end-to-end latency overhead.

FourTune: Towards Fully 4-Bit Efficient Post-Training for Diffusion Models

ICML 2026
 (
)

An end-to-end W4A4G4 post-training framework for diffusion models, cutting memory 2.25× and boosting training throughput 2.27× over BF16 LoRA on FLUX.1-dev.