Efficient AI Computing,
Transforming the Future.

Projects

To choose projects, simply check the boxes of the categories, topics and techniques.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

ICML 2024
 (
)

Quest is an effient long-context LLM inference framework that leverages query-aware sparsity in KV cache to reduce memory movement during attention and thus boost throughput.

Efficient Streaming Language Models with Attention Sinks

ICLR 2024
 (
)

We enable LLMs to work on infinite-length texts without compromising efficiency and performance.

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

MLSys 2024
 (
)

Low-bit weight-only quantization for LLMs.

LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

ICLR 2024
 (
)

LongLoRA takes advantage of shifted sparse attention to greatly reduce the finetuning cost of long context LLMs.