ArtiFixer: Enhancing and Extending 3D Reconstruction
Enhances and extends sparse 3D reconstructions with an auto‑regressive diffusion model that generates hundreds of consistent novel views in a single pass.
AI methods for reconstructing, generating and editing dynamic 3D assets and scenes.
We build systems for dynamic 3D content across reconstruction and generation, as well as editing and interaction. For reconstruction, we focus on feed‑forward and neural techniques to reconstruct simulation‑ready worlds from sensor data, improving fidelity, robustness, and scalability. For generation, we develop controllable methods to create high‑quality 3D/4D assets and dynamic scenes from text, images, and videos, leveraging camera‑aware video diffusion models, mesh‑aware LLMs, and efficient 3D priors. For interaction and editing, we develop techniques to enhance, relight and re‑condition captured and generated content — harmonizing renderings, synthesizing and removing weather, and correcting artifacts for photorealistic simulation.
Text/Image/Video‑to‑3D/4D • Controllability • Scene Generation • Neural Reconstruction • Simulation‑Ready Assets • Interactive Editing • Real-Time Methods
AI methods for reconstructing and/or generating simulation environments in which end-to-end policy models can be tested, evaluated, and trained in closed-loop.
Enhances and extends sparse 3D reconstructions with an auto‑regressive diffusion model that generates hundreds of consistent novel views in a single pass.
Uses a controllable video generator as a prior to recover persistent, dynamic Gaussian reconstructions from monocular video.
Applies one looped transformer block recurrently, turning refinement steps into an inference‑time compute knob that matches far larger feed‑forward models at a fraction of the parameters.
Distills the varying‑length key‑value scene representation into a fixed‑size MLP, removing the quadratic cost that limits offline feed‑forward reconstruction to few input images.
Controllable methods to synthesize high‑quality 3D/4D assets and dynamic scenes.
Scalable synthetic driving data: controllable, multi‑view, temporally consistent video and LiDAR generated from world foundation models.
Autoregressively generates action‑conditioned video in real time, driving closed‑loop simulation on top of NVIDIA Cosmos.
Guiding video generation with explicit 3D caches for precise camera control and 3D consistency.
Generating unbounded dynamic driving scenes with world‑guided video models.
AI methods for editing, enhancing and relighting 3D content across representations — from captured to generated — turning imperfect renderings into coherent, photorealistic environments for simulation, films and games.
Single‑step, temporally‑conditioned enhancer that harmonizes inserted objects and removes reconstruction artifacts online, on a single GPU.
Controllable addition and removal of weather and scene conditions in video, preserving physical and spatial coherence.