Efficient AI
Efficient AI
News
Blog
Publications
Light
Dark
Automatic
Blogs
Autoregressive Video Gen
A technical survey of autoregressive video generation: why causality enables streaming and KV-cache reuse, how few-step distillation makes it practical, and how recent systems generate longer videos while reducing training cost.
Shuai Yang
,
Yukang Chen
,
Luozhou Wang
,
Wei Huang
,
Weian Mao
,
Song Han
Jul 13, 2026
Pushing Intelligence to 4-bit
Four-bit floating point is moving from a storage-only compression trick to a primitive for training and inference across LLMs, diffusion, video generation, KV cache, and attention.
Wei Huang
,
Yukang Chen
,
Weian Mao
,
Luozhou Wang
,
Shuai Yang
,
Song Han
Jun 16, 2026
KV Cache Compression and Its Infra Problems
Why KV cache compression methods that work on paper fail in production — FlashAttention hides attention scores, and repeated token eviction frees no paged GPU memory — and how a geometric property hidden beneath RoPE resolves both.
Weian Mao
,
Yukang Chen
,
Wei Huang
,
Shuai Yang
,
Luozhou Wang
,
Song Han
Jun 12, 2026
Scaling Video Training with Parallelism
Long-video training changes the unit of distributed computation. This blog explains how sequence parallelism scales training when one video sample is too long for one GPU, comparing LongVILA MM-SP and LongLive-2.0 Balanced SP.
Yukang Chen
,
Luozhou Wang
,
Wei Huang
,
Shuai Yang
,
Weian Mao
,
Song Han
Jun 3, 2026
Why Video Gen Is an Infra Problem
Video generation is becoming an infrastructure problem: long context, memory, decoding, scheduling, compression, parallelism, and deployment-aware design.
Yukang Chen
,
Luozhou Wang
,
Wei Huang
,
Shuai Yang
,
Weian Mao
,
Song Han
May 26, 2026
Cite
×