OmniDreams
Real-time generative closed-loop autonomous vehicle simulation built on NVIDIA Cosmos.
Modeling dynamic worlds and the agents that perceive, act, and respond within them.
World simulation requires modeling how a scene looks and evolves, as well as how agents act within it. We study these connected problems through world modeling and action modeling. World modeling reconstructs and generates dynamic environments; action modeling develops policies for real-world agents and reactive behaviors for simulated ones. Together, they support interactive environments for developing and evaluating physical AI.
World generation • Dynamic scene reconstruction • Agent policies • Reactive simulation
We build generative world models that learn predictive, controllable representations of dynamic scenes across space and time. Our work combines video generation with explicit 3D/4D scene representations to reconstruct observed worlds, imagine plausible content beyond the available views, and maintain geometric and temporal consistency. These models support camera-controlled novel-view synthesis, dynamic-scene reconstruction, and scalable closed-loop simulation for training and evaluating intelligent agents.
Real-time generative closed-loop autonomous vehicle simulation built on NVIDIA Cosmos.
Enhances and extends sparse 3D reconstructions with an autoregressive diffusion model that generates consistent novel views.
Scalable synthetic driving data with controllable, multi‑view, temporally consistent videos and LiDAR from world models.
Uses a controllable video generator as a prior to recover persistent, dynamic Gaussian reconstructions from monocular video.
We study how agents understand dynamic scenes, choose actions, and respond to one another. The work spans two complementary settings: policies that operate in the real world and simulations that react to those policies. Compact visual representations can support both, while their goals differ—real-world policies choose ego actions, whereas simulations model how the scene and non-ego agents respond. We explore two ways to make those simulations reactive: generating appearance and behavior jointly, and driving agents with policies learned through multi-agent reinforcement learning. Policy auto-research then provides a way for coding agents to autonomously improve those policies.
From visual inputs such as multi-camera video streams, we learn compact representations for action prediction and world simulation. To build such representations, we reason jointly across cameras and timestamps, stripping away redundancy while preserving the crucial information downstream tasks need. By capturing the essential state of a scene, these representations inform both action prediction—“What to do?”—and world simulation—“What happens next?”
Most generative world models produce photorealistic pixels without explicitly modeling agent behavior. As a result, other agents in the scene cannot react to novel actions taken by the ego agent. We study models that jointly generate world appearance and agent behavior, enabling non-ego agents to respond realistically to novel ego actions.
Multi-agent reinforcement learning (MARL) produces reasonable policies by rewarding each agent for reaching its goal while avoiding collisions. These policies drive the non-ego agents surrounding the ego agent, creating reactive simulation environments for training and evaluating physical AI models. Simulation is most valuable when it exposes models to rare, long-tail behaviors. We therefore study how to steer MARL toward desired behaviors, including by conditioning policies on natural-language prompts.
Coding agents can now execute long-running tasks autonomously. Can they also make policies that control non-ego agents, making the environment responsive to novel ego actions? We explore an automated research framework in which coding agents propose hypotheses, evaluate them systematically, summarize what they learn, use those findings to plan the next round of experiments, and repeat the cycle until they produce a satisfactory policy.