Videos generated with PDD-LTX-2.3 and PDD-Wan2.1-14B models. Click the top right of LTX-2.3 videos for sound.
Playback quality:

8 NFE PDD on LTX-2.3 with upscaler & refiner: The video depicts a grand scene where a king, adorned in ornate armor, stands at the center, flanked by a group of soldiers also clad in elaborate, dark blue armor with gold accents, the heavy metallic clinking of their gear echoing softly in the cavernous, stone-walled chamber. Each soldier holds a sword, their blades gleaming under the dimly lit environment, suggesting an impending battle or ceremonial display as the faint, rhythmic rasp of steel sliding against leather scabbards punctuates the heavy silence. The soldiers' formation is disciplined, with some standing upright while others appear slightly bent forward, creating a sense of readiness and anticipation amidst the low, steady hum of distant torches flickering in the drafty air. The king's commanding presence is emphasized by his elevated position and the respectful stance of the surrounding soldiers, and as he surveys his ranks, he speaks in a deep, gravelly tone, his voice resonating with authority through the hall, "Stand fast, for today we carve our legacy." The overall atmosphere conveys a mix of tension and solemnity, with the muffled, ambient weight of the stone architecture grounding the significance of the event unfolding.
4 NFE PDD on Wan2.1 14B: Two anthropomorphic cats, one with sleek black fur and the other with fluffy white fur, are engaged in an intense boxing match on a spotlighted stage. They are dressed in vibrant, comfy boxing gear with matching bright red gloves. The black-furred cat has a determined look on its face, while the white-furred cat appears focused and ready to strike. Both cats move dynamically, throwing punches and dodging each other with agility. The background showcases a bustling audience cheering them on, with colorful lights and banners creating an energetic atmosphere. Medium shot capturing the dynamic action from a side angle, emphasizing the motion and interaction between the two fighters.
4 NFE PDD on Wan2.1 14B: A majestic large bird, with broad wings spread out, flying gracefully towards an ancient church tower adorned with intricate gothic architecture. The large ticking clock mounted on the tower shows 6 AM, casting a soft glow over the scene. The bird's feathers shimmer in the early morning light as it soars through the air. The background reveals a serene, quiet village with early morning mist hovering over it. The shot starts from a mid-range view of the bird in flight, then smoothly transitions to a closer view focusing on the bird as it approaches the towering church. The camera captures the detailed movement of the bird's wings and the rhythmic ticking of the clock, creating a tranquil and atmospheric scene.
4 NFE PDD on Wan2.1 14B: A joyful child, with a big smile and arms spread wide, swings energetically on a rusty old swing set in a sunlit backyard. The swing set, with peeling paint and creaking chains, contrasts against the vibrant green grass and blooming flowers surrounding it. The child's laughter echoes as they swing higher and higher, their feet barely touching the ground at the bottom of each arc. The scene is captured from a low angle, emphasizing the height of the swings, with the sun casting a warm glow over everything. Medium shot focusing on the child and the swing set.
4 NFE PDD on Wan2.1 14B: Two cyclists, one male and one female, are approaching a stop sign at the intersection from opposite directions. The male cyclist is wearing a red helmet and black cycling gear, while the female cyclist is wearing a blue helmet and white cycling clothes. They are both pedaling towards the stop sign, gradually slowing down as they near the intersection. The male cyclist is glancing left, checking for traffic, while the female cyclist is focused straight ahead. The background shows a quiet suburban street with green trees and parked cars. The video captures their approach in a tracking shot, gradually zooming in on each cyclist as they slow to a stop, maintaining eye-level perspectives.
8 NFE PDD on LTX-2.3 with upscaler & refiner: A smiling young woman stands confidently in a dense forest, the air filled with the gentle, rhythmic rustling of leaves and the distant, melodic chirping of songbirds echoing through the open woodland. She has fair skin, green eyes, and long wavy blonde hair flowing freely, catching the light as she takes a deep, audible breath of fresh air. She is dressed in a casual outfit consisting of a light green blouse and olive-green cargo pants, the fabric softly swishing as her arms gently swing at her sides. The forest is filled with tall trees and dappled sunlight filtering through the canopy, while her footsteps make a soft, muffled crunch on the moss-covered ground. The camera captures her from a medium close-up angle, focusing on her joyful expression as she lets out a contented, airy sigh, the serene environment alive with the peaceful, immersive sound of a breeze moving through the branches.
8 NFE PDD on LTX-2.3 with upscaler & refiner: A person with a distressed and angry expression shouts at another individual standing a few feet away, their voice cracking with raw intensity as they yell, "I can't believe you did this!" while leaning forward with arms outstretched and fists clenched, the sound of their heavy, ragged breathing filling the room. The second person stands in silence, arms crossed defensively and eyes cast downward to avoid contact, the only sound being the faint, rhythmic creak of a wooden chair in the dimly lit room. The camera captures the intense emotion in a medium close-up shot, the acoustic space feeling cramped and stifling as the first person glares, their voice dropping to a harsh, trembling mutter, "Just leave," followed by the sharp, echoing thud of a fist hitting the wooden table.
4 NFE PDD on Wan2.1 14B: In the style of Vincent van Gogh, an astronaut with a helmet and spacesuit rides a large cow on a sandy beach at sunset. The astronaut sits confidently on the cow, which is walking steadily towards the viewer. The cow has a serene expression and a gentle gait. To the left of the astronaut and the cow stands a tall, swaying palm tree with vibrant, swirling leaves. The sky behind them is filled with vivid, swirling strokes of orange, yellow, and blue, capturing the essence of a Van Gogh starry night. Medium shot, focusing on the astronaut and the cow, with the palm tree and beach visible in the background.
4 NFE PDD on Wan2.1 14B: In an energetic live concert setting, a skilled guitarist stands center stage under bright spotlights. He plays an electric guitar with intense focus and passion. His fingers swiftly move across the fretboard as he strums the strings with vigor. The camera starts from a wide shot capturing the dynamic atmosphere of the packed venue before slowly zooming in on the guitarist's hands and the guitar, highlighting the intricate movements and the vibrant glow of the stage lights reflecting off the instrument. Close-up shot focusing on the guitar and the musician's hands.

Abstract

Figure 1: Parallel Decoding Distillation (PDD) vs. Flow Matching (FM): with each network evaluation, PDD advances multiple steps, while FM advances only one.

Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generators. Although these methods achieve high-quality video generation, their training losses are notoriously hard to optimize and suffer from mode collapse, leading to a lack of video diversity and motion.

In this paper, we introduce Parallel Decoding Distillation (PDD), a simplified and scalable trajectory-based distillation method for fast inference of diffusion and flow matching models. Our architecture and training procedure are compatible with any pre-trained model and support sampling with a varying number of function evaluations (NFE). PDD accelerates generation by predicting multiple denoising steps per network evaluation. Conceptually, it learns a representation of the mean velocity without regressing its derivative using JVPs or finite-difference approximations. Instead, it decomposes the mean-velocity prediction over a fixed interval into parallel sub-interval predictions. Then, optimizing over a single sub-interval per training iteration provides the signal for learning the full-interval mean velocity.

Our method achieves SOTA performance with 4-8 NFE on LTX-2.3 Text-to-Video/Audio, Wan2.1 14B Text-to-Video, and Qwen-Image Text-to-Image. Moreover, PDD shows a significant improvement in generated video diversity.

Parallel Decoding Distillation (PDD): Intuition

PDD sketch: the sampling trajectory is discretized into N intervals grouped into blocks of size L; the parallel decoder predicts all mean velocities in a block in one evaluation.
Figure 2: The sampling trajectory is discretized into \(N\) intervals, which are grouped into blocks of size \(L\). The parallel decoder predicts the mean velocities for all intervals within a block using a single evaluation.

Sampling diffusion and flow matching models using classical ordinary differential equation (ODE) solvers requires many small denoising steps, resulting in hundreds of network evaluations. The common strategy of trajectory-based distillation methods for training a few-step generator is to train a student model to predict a large-step update, essentially merging multiple small teacher updates. In Parallel Decoding Distillation (PDD), we take a different approach: instead of merging the teacher updates into one large update, we predict multiple teacher updates in a single network evaluation. Specifically, we discretize time \(t\in[0,1]\) into a fixed sequence of \(N\in\mathbb{N}\) intervals \([t_{n},t_{n+1}]\) for \(n=0,1,\ldots,N-1\) and group them into blocks of size \(L\leq N\). During training, the student model learns to predict the teacher's ODE-solver step for all intervals within a block using a single forward pass. Then, at generation, each student evaluation advances \(L\) intervals, as illustrated in Figure 2.

Teacher: the pre-trained flow model architecture.
(a) Teacher
Student training: the parallel decoder architecture.
(b) Student training
Student generation: the fused parallel decoder architecture.
(c) Student Generation
Figure 3: Architecture of the parallel decoder (b) vs. the pre-trained flow model (a). Notably, the parallel decoder utilizes the same backbone, but with the final linear layer replicated \(N\) times. (c) At generation, instead of applying a linear layer per intra-block step, we can fuse the layers into a single linear layer that directly predicts a step across the full block.

Figure 3(b) illustrates our student architecture, alongside the teacher in Figure 3(a). We start from the teacher architecture and simply repeat the final linear layer \(N\) times, where \(N\) is the grid size. That is, we use a distinct final linear layer for each interval \([t_n,t_{n+1}]\) in the discretization. Predicting \(N\) grid-level outputs, rather than only \(L\) block-level outputs, enables sampling with varying NFEs at inference time. We achieve this by training with variable block sizes and then adjusting the block size at generation according to the desired NFE. Moreover, since a weighted sum of linear transformations is itself a linear transformation, we can fuse the relevant linear layers at generation. As a result, for each block, we only need to evaluate and store a single fused linear layer; see Figure 3(c). Lastly, this architecture allows the student weights to be initialized seamlessly from the teacher weights.

Student and teacher trajectories during training.
(a) Student and teacher during training
Parallel decoding loss sketch.
(b) Parallel decoding loss sketch
Figure 4: Blue vectors depict the student rollout, the green vector the ODE-solver step over the interval \([t_n, t_{n+1}]\), and the red vector the residual penalized by the parallel decoding loss.

To train our parallel decoder (student), we employ an on-policy training algorithm. The goal is to align each student output with the corresponding teacher ODE-solver step. That is, we align the \(n\)-th student output with the teacher ODE-solver step over the interval \([t_n, t_{n+1}]\). We illustrate a single optimization step with block size \(L=4\) in Figure 4. First, we sample a noisy state \(X_n\) at time step \(n\). Then, starting from \(X_n\), we roll out the student over the next block of intervals:

\[ X_n=\bar{X}_n,\ \bar{X}_{n+1},\ldots,\bar{X}_{n+L-1}. \]

Next, we sample a single state \(\bar{X}_k\) from the student trajectory, where \(k\in\{n,n+1,\ldots,n+L-1\}\). Finally, we match the \(k\)-th student output to the teacher's ODE-solver step evaluated at \(\bar{X}_k\). Figure 5 shows a simulation of PDD training on a toy problem. For each iteration, we plot the largest-error sample from the batch. Notably, the student trajectory becomes aligned with the teacher trajectory (labeled “Optimal trajectory”) as training progresses.

Figure 5: (Left) Animation of PDD training across multiple iterations, showing one batch sample at each iteration. (Middle) Zoom-in on the sampled state along the student trajectory, comparing the student output with the teacher output at that state. (Right) PDD loss over training iterations.

PDD vs. Teacher and DMD2 on Wan2.1 14B T2V

Comparison between the teacher (Wan2.1 14B T2V), PDD, and DMD2. The teacher requires 2 network evaluations per step due to CFG. Each row is a method; the three videos are three different noise samples for the same prompt. Use the side arrows or swipe to browse all 23 prompts.

PDD vs. AnyFlow on Wan2.1 14B T2V

Comparison between PDD and AnyFlow on Wan2.1 14B T2V. Each method is one row; the first two videos are 4 NFE and the last two are 8 NFE, for the same prompt. Use the side arrows or swipe to browse all 22 prompts.

PDD vs. Teacher and official distilled model on LTX-2.3

Comparison between the teacher (LTX-2.3), PDD, and the official distilled model. We use the same official upscaler and refiner for all three methods. The teacher requires 4 network evaluations per step due to guidance. Each row is a method. Click the top right of each video for sound, and use the side arrows or swipe to browse all 12 prompts.