HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers
Diffusion-based video relighting enables controllable relighting from a single input video, but modern video diffusion backbones are trained on short clips and applied to long-horizon videos through chunked sliding-window inference, often causing temporal discontinuities at chunk boundaries. We address this by treating long-horizon relighting as a temporally conditioned latent-domain translation problem. Our framework enforces cross-chunk continuity by propagating target-domain latents across boundaries. Masked self-conditioning in the target domain teaches the model to continue from temporally masked propagated context. We further introduce warm-start prompting, which uses a relit prompt anchor from a controllable generative model to establish the initial target-domain state and provide a general interface for prompt-based relighting. Experiments on in-the-wild long-horizon videos show markedly improved temporal consistency, substantially reducing chunk-boundary artifacts and suppressing unwanted appearance changes between chunks.