Efficient AI
Efficient AI
News
Blog
Publications
Light
Dark
Automatic
KV Cache
Autoregressive Video Gen
A technical survey of autoregressive video generation: why causality enables streaming and KV-cache reuse, how few-step distillation makes it practical, and how recent systems generate longer videos while reducing training cost.
Shuai Yang
,
Yukang Chen
,
Luozhou Wang
,
Wei Huang
,
Weian Mao
,
Song Han
Jul 13, 2026
KV Cache Compression and Its Infra Problems
Why KV cache compression methods that work on paper fail in production — FlashAttention hides attention scores, and repeated token eviction frees no paged GPU memory — and how a geometric property hidden beneath RoPE resolves both.
Weian Mao
,
Yukang Chen
,
Wei Huang
,
Shuai Yang
,
Luozhou Wang
,
Song Han
Jun 12, 2026
Cite
×