Conclusion


We challenge the prevailing view of TTT as test-time memorization. Through systematic empirical analysis and mathematical derivation, we show that TTT—even with complex inner loops involving multi-layer MLPs and momentum—is fundamentally a form of linear attention. This reframing is not merely theoretical: it enables principled simplifications that often improve performance, parallel implementations with up to 4× speedups, and a unified framework for understanding diverse TTT variants.