HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers

Diffusion-based video relighting enables controllable relighting from a single input video, but modern video diffusion backbones are trained on short clips and applied to long-horizon videos through chunked sliding-window inference, often causing temporal discontinuities at chunk boundaries. We address this by treating long-horizon relighting as a temporally conditioned latent-domain translation problem. Our framework enforces cross-chunk continuity by propagating target-domain latents across boundaries. Masked self-conditioning in the target domain teaches the model to continue from temporally masked propagated context. We further introduce warm-start prompting, which uses a relit prompt anchor from a controllable generative model to establish the initial target-domain state and provide a general interface for prompt-based relighting. Experiments on in-the-wild long-horizon videos show markedly improved temporal consistency, substantially reducing chunk-boundary artifacts and suppressing unwanted appearance changes between chunks.

Authors

Jing Yang (NVIDIA, University of Southern California)
Mayoore Jaiswal (NVIDIA)
Zian Wang (NVIDIA)
Xiao Steven Zeng (NVIDIA)
Yajie Zhao (University of Southern California)
Rochelle Pereira (NVIDIA)
Jianyuan Min (NVIDIA)

Publication Date