<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Enze Xie | Efficient AI</title><link>https://research.nvidia.com/labs/eai/author/enze-xie/</link><atom:link href="https://research.nvidia.com/labs/eai/author/enze-xie/index.xml" rel="self" type="application/rss+xml"/><description>Enze Xie</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Mon, 13 Oct 2025 00:00:00 +0000</lastBuildDate><image><url>https://research.nvidia.com/labs/eai/author/enze-xie/avatar_hu_3020942bad90eae6.png</url><title>Enze Xie</title><link>https://research.nvidia.com/labs/eai/author/enze-xie/</link></image><item><title>LongLive: Real-time Interactive Long Video Generation</title><link>https://research.nvidia.com/labs/eai/publication/longlive/</link><pubDate>Mon, 13 Oct 2025 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/longlive/</guid><description>&lt;!--
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Click the &lt;em&gt;Cite&lt;/em&gt; button above to demo the feature to enable visitors to import publication metadata into their reference management software.
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Create your slides in Markdown - click the &lt;em&gt;Slides&lt;/em&gt; button to check out the example.
&lt;/div&gt;
&lt;/div&gt;
Add the publication's **full text** or **supplementary notes** here. You can use rich formatting such as including [code, math, and images](https://docs.hugoblox.com/content/writing-markdown-latex/). --&gt;</description></item><item><title>Fast-dLLM v2: Efficient Block-Diffusion LLM</title><link>https://research.nvidia.com/labs/eai/publication/fast-dllm-v2/</link><pubDate>Tue, 30 Sep 2025 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/fast-dllm-v2/</guid><description>&lt;!--
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Click the &lt;em&gt;Cite&lt;/em&gt; button above to demo the feature to enable visitors to import publication metadata into their reference management software.
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Create your slides in Markdown - click the &lt;em&gt;Slides&lt;/em&gt; button to check out the example.
&lt;/div&gt;
&lt;/div&gt;
Add the publication's **full text** or **supplementary notes** here. You can use rich formatting such as including [code, math, and images](https://docs.hugoblox.com/content/writing-markdown-latex/). --&gt;</description></item><item><title>DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space</title><link>https://research.nvidia.com/labs/eai/publication/dc-gen/</link><pubDate>Mon, 29 Sep 2025 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/dc-gen/</guid><description>&lt;!--
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Click the &lt;em&gt;Cite&lt;/em&gt; button above to demo the feature to enable visitors to import publication metadata into their reference management software.
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Create your slides in Markdown - click the &lt;em&gt;Slides&lt;/em&gt; button to check out the example.
&lt;/div&gt;
&lt;/div&gt;
Add the publication's **full text** or **supplementary notes** here. You can use rich formatting such as including [code, math, and images](https://docs.hugoblox.com/content/writing-markdown-latex/). --&gt;</description></item><item><title>DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder</title><link>https://research.nvidia.com/labs/eai/publication/dc-videogen/</link><pubDate>Mon, 29 Sep 2025 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/dc-videogen/</guid><description>&lt;!--
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Click the &lt;em&gt;Cite&lt;/em&gt; button above to demo the feature to enable visitors to import publication metadata into their reference management software.
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Create your slides in Markdown - click the &lt;em&gt;Slides&lt;/em&gt; button to check out the example.
&lt;/div&gt;
&lt;/div&gt;
Add the publication's **full text** or **supplementary notes** here. You can use rich formatting such as including [code, math, and images](https://docs.hugoblox.com/content/writing-markdown-latex/). --&gt;</description></item><item><title>SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer</title><link>https://research.nvidia.com/labs/eai/publication/sana-video/</link><pubDate>Mon, 29 Sep 2025 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/sana-video/</guid><description>&lt;!--
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Click the &lt;em&gt;Cite&lt;/em&gt; button above to demo the feature to enable visitors to import publication metadata into their reference management software.
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Create your slides in Markdown - click the &lt;em&gt;Slides&lt;/em&gt; button to check out the example.
&lt;/div&gt;
&lt;/div&gt;
Add the publication's **full text** or **supplementary notes** here. You can use rich formatting such as including [code, math, and images](https://docs.hugoblox.com/content/writing-markdown-latex/). --&gt;</description></item><item><title>DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space</title><link>https://research.nvidia.com/labs/eai/publication/dcae-1.5/</link><pubDate>Fri, 01 Aug 2025 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/dcae-1.5/</guid><description>&lt;!--
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Click the &lt;em&gt;Cite&lt;/em&gt; button above to demo the feature to enable visitors to import publication metadata into their reference management software.
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Create your slides in Markdown - click the &lt;em&gt;Slides&lt;/em&gt; button to check out the example.
&lt;/div&gt;
&lt;/div&gt;
Add the publication's **full text** or **supplementary notes** here. You can use rich formatting such as including [code, math, and images](https://docs.hugoblox.com/content/writing-markdown-latex/). --&gt;</description></item><item><title>DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer</title><link>https://research.nvidia.com/labs/eai/publication/dc-ar/</link><pubDate>Mon, 07 Jul 2025 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/dc-ar/</guid><description>&lt;!--
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Click the &lt;em&gt;Cite&lt;/em&gt; button above to demo the feature to enable visitors to import publication metadata into their reference management software.
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Create your slides in Markdown - click the &lt;em&gt;Slides&lt;/em&gt; button to check out the example.
&lt;/div&gt;
&lt;/div&gt;
Add the publication's **full text** or **supplementary notes** here. You can use rich formatting such as including [code, math, and images](https://docs.hugoblox.com/content/writing-markdown-latex/). --&gt;</description></item><item><title>Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding</title><link>https://research.nvidia.com/labs/eai/publication/fast-dllm/</link><pubDate>Wed, 28 May 2025 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/fast-dllm/</guid><description>&lt;!--
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Click the &lt;em&gt;Cite&lt;/em&gt; button above to demo the feature to enable visitors to import publication metadata into their reference management software.
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Create your slides in Markdown - click the &lt;em&gt;Slides&lt;/em&gt; button to check out the example.
&lt;/div&gt;
&lt;/div&gt;
Add the publication's **full text** or **supplementary notes** here. You can use rich formatting such as including [code, math, and images](https://docs.hugoblox.com/content/writing-markdown-latex/). --&gt;</description></item><item><title>SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation</title><link>https://research.nvidia.com/labs/eai/publication/sana-sprint/</link><pubDate>Wed, 12 Mar 2025 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/sana-sprint/</guid><description>&lt;!--
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Click the &lt;em&gt;Cite&lt;/em&gt; button above to demo the feature to enable visitors to import publication metadata into their reference management software.
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Create your slides in Markdown - click the &lt;em&gt;Slides&lt;/em&gt; button to check out the example.
&lt;/div&gt;
&lt;/div&gt;
Add the publication's **full text** or **supplementary notes** here. You can use rich formatting such as including [code, math, and images](https://docs.hugoblox.com/content/writing-markdown-latex/). --&gt;</description></item><item><title>VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation</title><link>https://research.nvidia.com/labs/eai/publication/vila-u/</link><pubDate>Tue, 04 Mar 2025 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/vila-u/</guid><description/></item><item><title>SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer</title><link>https://research.nvidia.com/labs/eai/publication/sana-1.5/</link><pubDate>Mon, 20 Jan 2025 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/sana-1.5/</guid><description>&lt;!--
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Click the &lt;em&gt;Cite&lt;/em&gt; button above to demo the feature to enable visitors to import publication metadata into their reference management software.
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Create your slides in Markdown - click the &lt;em&gt;Slides&lt;/em&gt; button to check out the example.
&lt;/div&gt;
&lt;/div&gt;
Add the publication's **full text** or **supplementary notes** here. You can use rich formatting such as including [code, math, and images](https://docs.hugoblox.com/content/writing-markdown-latex/). --&gt;</description></item><item><title>SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models</title><link>https://research.nvidia.com/labs/eai/publication/svdquant/</link><pubDate>Thu, 07 Nov 2024 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/svdquant/</guid><description>&lt;h3 id="overview"&gt;Overview&lt;/h3&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="d-flex justify-content-center"&gt;
&lt;div class="w-100" &gt;&lt;img src="https://github.com/mit-han-lab/nunchaku/blob/main/assets/demo.gif?raw=true" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
SVDQuant is a post-training quantization technique for 4-bit weights and activations that well maintains visual fidelity. On 12B FLUX.1-dev, it achieves 3.6× memory reduction compared to the BF16 model. By eliminating CPU offloading, it offers 8.7× speedup over the 16-bit model when on a 16GB laptop 4090 GPU, 3× faster than the NF4 W4A16 baseline. On PixArt-∑, it demonstrates significantly superior visual quality over other W4A4 or even W4A8 baselines. &amp;ldquo;E2E&amp;rdquo; means the end-to-end latency including the text encoder and VAE decoder.&lt;/p&gt;
&lt;h3 id="method"&gt;Method&lt;/h3&gt;
&lt;h4 id="svdquant-absorbing-outliers-via-low-rank-branch"&gt;SVDQuant: Absorbing Outliers via Low-Rank Branch&lt;/h4&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="d-flex justify-content-center"&gt;
&lt;div class="w-100" &gt;&lt;img src="https://github.com/mit-han-lab/nunchaku/blob/main/assets/intuition.gif?raw=true" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;Stage 1: Originally, both the activation $\boldsymbol{X}$
and weights $\boldsymbol{W}$
contain outliers, making 4-bit quantization challenging. Stage 2: We migrate the outliers from activations to weights, resulting in the updated activation $\hat{\boldsymbol{X}}$
and weights $\hat{\boldsymbol{W}}$
. While $\hat{\boldsymbol{X}}$
becomes easier to quantize, $\hat{\boldsymbol{W}}$
now becomes more difficult. Stage 3: SVDQuant further decomposes $\hat{\boldsymbol{W}}$
into a low-rank component $\boldsymbol{L}_1 \boldsymbol{L}_2$
and a residual $\hat{\boldsymbol{W}} - \boldsymbol{L}_1 \boldsymbol{L}_2$
with SVD. Thus, the quantization difficulty is alleviated by the low-rank branch, which runs at 16-bit precision.&lt;/p&gt;
&lt;h4 id="nunchaku-fusing-low-rank-and-low-bit-branch-kernels"&gt;Nunchaku: Fusing Low-Rank and Low-Bit Branch Kernels&lt;/h4&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="d-flex justify-content-center"&gt;
&lt;div class="w-100" &gt;&lt;img src="https://cdn.prod.website-files.com/64f4e81394e25710d22d042e/672d1fccad9d41d739c84c2c_672a6ac09cf7fc46071c3fb8_672a6a991adf3346a0ee4dbb_engine.jpeg" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;(a) Naïvely running low-rank branch with rank 32 will introduce 57% latency overhead due to extra read of 16-bit inputs in Down Projection and extra write of 16-bit outputs in Up Projection. Nunchaku optimizes this overhead with kernel fusion. (b) Down Projection and Quantize kernels use the same input, while Up Projection and 4-Bit Compute kernels share the same output. To reduce data movement overhead, we fuse the first two and the latter two kernels together.&lt;/p&gt;
&lt;h3 id="performance"&gt;Performance&lt;/h3&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="d-flex justify-content-center"&gt;
&lt;div class="w-100" &gt;&lt;img src="https://cdn.prod.website-files.com/64f4e81394e25710d22d042e/672d1bcdf3c3ec127e8078f3_672d1b5772f1247fa24e231d_efficiency.jpeg" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
SVDQuant reduces the model size of the 12B FLUX.1 by 3.6×. Additionally, Nunchaku further cuts memory usage of the 16-bit model by 3.5× and delivers 3.0× speedups over the NF4 W4A16 baseline on both the desktop and laptop NVIDIA RTX 4090 GPUs. Remarkably, on laptop 4090, it achieves in total 10.1× speedup by eliminating CPU offloading.&lt;/p&gt;
&lt;h3 id="integrate-with-lora"&gt;Integrate with LoRA&lt;/h3&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="d-flex justify-content-center"&gt;
&lt;div class="w-100" &gt;&lt;img src="https://github.com/mit-han-lab/nunchaku/blob/main/assets/lora.jpg?raw=true" alt="" loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
SVDQuant seamlessly integrates with off-the-shelf LoRAs without requiring re-quantization. When applying LoRAs, it matches the image quality of the original 16-bit FLUX.1-dev.&lt;/p&gt;
&lt;h3 id="video"&gt;Video&lt;/h3&gt;
&lt;iframe width="720" height="450" src="https://www.youtube.com/embed/nYujDH9r69s?si=d-6cMb9HU-4szCmQ" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen&gt;&lt;/iframe&gt;
&lt;h5 id="acknowledgment"&gt;Acknowledgment&lt;/h5&gt;
&lt;p&gt;We thank MIT-IBM Watson AI Lab, MIT and Amazon Science Hub, MIT AI Hardware Program, National Science Foundation, Packard Foundation, Dell, LG, Hyundai, and Samsung for supporting this research. We thank NVIDIA for donating the DGX server.&lt;/p&gt;</description></item><item><title>Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models</title><link>https://research.nvidia.com/labs/eai/publication/dcae/</link><pubDate>Mon, 14 Oct 2024 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/dcae/</guid><description>&lt;!--
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Click the &lt;em&gt;Cite&lt;/em&gt; button above to demo the feature to enable visitors to import publication metadata into their reference management software.
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="alert alert-note"&gt;
&lt;div&gt;
Create your slides in Markdown - click the &lt;em&gt;Slides&lt;/em&gt; button to check out the example.
&lt;/div&gt;
&lt;/div&gt;
Add the publication's **full text** or **supplementary notes** here. You can use rich formatting such as including [code, math, and images](https://docs.hugoblox.com/content/writing-markdown-latex/). --&gt;</description></item><item><title>HART: Efficient Visual Generation with Hybrid Autoregressive Transformer</title><link>https://research.nvidia.com/labs/eai/publication/hart/</link><pubDate>Mon, 14 Oct 2024 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/hart/</guid><description/></item><item><title>SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer</title><link>https://research.nvidia.com/labs/eai/publication/sana/</link><pubDate>Mon, 14 Oct 2024 00:00:00 +0000</pubDate><guid>https://research.nvidia.com/labs/eai/publication/sana/</guid><description>&lt;h3 id="oral-presentation-video"&gt;Oral Presentation Video&lt;/h3&gt;
&lt;iframe width="100%" style="aspect-ratio: 1920 / 820;" src="https://www.youtube.com/embed/rrKFyYx19UI?si=cSrnlS5WOXpCDl3K" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen&gt;&lt;/iframe&gt;</description></item></channel></rss>