1. [Publications](/publications)
2. Hymba: A Hybrid-head Architecture for Small Language Models
 
 # Hymba: A Hybrid-head Architecture for Small Language Models

  ![Publication image](/sites/default/files/styles/wide/public/default_images/default.jpeg?itok=qUFsuJCP "Publication image")

 We propose Hymba, a family of small language models featuring a hybrid-head parallel architecture that integrates attention mechanisms and state space models (SSMs) within the same layer, offering parallel and complementary processing of the same inputs. In this hybrid-head module, attention heads provide high-resolution recall, while SSM heads facilitate efficient context summarization. Additionally, we introduce learnable meta tokens, which are prepended to prompts to store critical meta information, guiding subsequent tokens and alleviating the “forced-to-attend” burden associated with attention mechanisms. Thanks to the global context summarized by SSMs, the attention heads in our model can be further optimized through cross-layer key-value (KV) sharing and a mix of global and local attention, resulting in a compact cache size without compromising accuracy. Notably, Hymba achieves state-of-the-art performance among small LMs: Our Hymba-1.5B-Base model surpasses all sub-2B public models and even outperforms Llama-3.2-3B, achieving 1.32% higher average accuracy, an 11.67 reduction in cache size, and 3.49 higher throughput.



 ## Authors



Xin Dong* (NVIDIA)

[Yonggan Fu\*](/person/yonggan-fu)

Shizhe Diao (NVIDIA)

[Wonmin Byeon](/person/wonmin-byeon)

Zijia Chen (NVIDIA)

Ameya Sunil Mahabaleshwarkar (NVIDIA)

Shih-Yang Liu (NVIDIA, HKUST)

[Matthijs Van keirsbilck](/person/matthijs-van-keirsbilck)

[Min-Hung Chen](/person/min-hung-chen)

[Yoshi Nishi](/person/yoshi-nishi)

Yingyan Celine Lin (NVIDIA, Georgia Tech)

[Jan Kautz](/person/jan-kautz)

[Pavlo Molchanov](/person/pavlo-molchanov)

 

 

 ## Publication Date



Thursday, April 24, 2025

 

 ## Published in



[Hymba - ICLR 2025](https://jankautz.com/publications/Hymba_ICLR25.pdf)

 

 ## Research Area



[Artificial Intelligence and Machine Learning ](/research-area/machine-learning-artificial-intelligence)

[Natural Language Processing](/research-area/natural-language-processing)

 

 

 ## External Links



[openreview](https://openreview.net/forum?id=A1ztozypga)

 

 

 ## Uploaded Files



[Hymba\_ICLR25.pdf](https://d1qx31qr3h6wln.cloudfront.net/publications/Hymba_ICLR25.pdf "Open file in new window")6 MB

 

 

 ## Award



ICLR spotlight paper