Efficient AI
Efficient AI
News
Blog
Publications
Light
Dark
Automatic
LLM Inference
Pushing Intelligence to 4-bit
Four-bit floating point is moving from a storage-only compression trick to a primitive for training and inference across LLMs, diffusion, video generation, KV cache, and attention.
Wei Huang
,
Yukang Chen
,
Weian Mao
,
Luozhou Wang
,
Shuai Yang
,
Song Han
Jun 16, 2026
KV Cache Compression and Its Infra Problems
Why KV cache compression methods that work on paper fail in production — FlashAttention hides attention scores, and repeated token eviction frees no paged GPU memory — and how a geometric property hidden beneath RoPE resolves both.
Weian Mao
,
Yukang Chen
,
Wei Huang
,
Shuai Yang
,
Luozhou Wang
,
Song Han
Jun 12, 2026
Cite
×