Grounded-Exo2Ego: Structured Semantic Grounding for Robust
Exocentric-to-Egocentric Video Generation

NVIDIA

TLDR: Exo-to-ego is an important topic that can help solve robots' need for in-the-wild ego videos and enable new VR experiences. We propose a systematic solution at both the data and architecture level. We also go beyond the popular geometry-conditioned view synthesis approach, and show that semantic grounding helps further improve quality.

Exo input
Vista4D
EgoX
Ours
GT
Exo input
Vista4D
EgoX
Ours
GT
Exo input
Vista4D
EgoX
Ours
GT
Exo input
Vista4D
EgoX
Ours
GT

Auto-playing clips, same exo input fed through every pipeline. CPR and Cooking are New Activities; Omelet and Bike repair are Unseen. Full New Activities + Unseen comparisons appear in the "Qualitative comparisons" section below.

If the videos don't play, please click the video boxes or change to a different browser.

Abstract

Generating egocentric video from a single exocentric video is an emerging and important topic for AR/VR and physical AI. Compared with conventional novel view synthesis, exo-to-ego generation is a significantly harder task because the standard geometric conditioning becomes highly unreliable under extreme view changes and large unobservable regions.

We present Grounded-Exo2Ego, a principled framework that addresses these challenges at both the architectural and data levels. Architecturally, Grounded-Exo2Ego is a dual-branch video diffusion model that couples a geometric anchoring branch, which conditions the generation on the rendering of a 3D reconstruction, with a novel semantic grounding branch, which goes beyond the prevailing geometry-based approach and improves quality by synthesizing challenging regions based on object-level context.

Additionally, we found that the overlooked issue of camera-reconstruction misalignment severely undermines exo-to-ego learning. We thus introduce a camera re-localization algorithm that resolves this issue and substantially improves quality across all metrics. We further develop a fully automated synthetic data engine that generates and renders rigged 3D characters in procedurally generated environments. Evaluation on the challenging EgoExo4D dataset shows that our method outperforms recent state-of-the-art approaches by large margins across all metrics. Detailed ablations validate improvements from each of our contributions at both the data and architectural level.

Grounded-Exo2Ego teaser figure

Grounded-Exo2Ego. From a single exocentric video (top left), we lift the scene into a semantically grounded 3D reconstruction that anchors a dual-branch video diffusion model. The generated egocentric views (bottom) are substantially more accurate than those from existing methods.

Exo → Ego inherently harder than general view synthesis

(a) Exo input (a) Exo input
(b) 3D at ego (b) 3D at ego
(c) Vista4D at ego (c) Vista4D at ego
(d) EgoX (d) EgoX
(e) Ours (e) Ours
(f) GT (f) GT

SOTA (e.g. Vista4D) often do monocular depth estimation from the input view (a) and render from the ego view (b) to generate the conditioning ego rendering video. However, the ego rendering (b) is often riddled with holes and corrupted geometry, leading to bad results (c). Recent exo-to-ego methods such as EgoX perform better (d) but fail to capture the exact viewpoint and content. Ours (e) produces ego videos much closer to the ground truth (f).

Method overview

Method overview
Method overview. Geometric anchoring (top): Exocentric color and depth are lifted to a 3D reconstruction and rendered at the ego pose. Semantic grounding (bottom): Per-object context is extracted as segmentation masks and object tokens. The segmentation masks are reprojected into the ego-view as used as attention masks during cross-attention. The per-object masks encourages noisy tokens to attend to object information relevant to the specific region, providing spatial structure that grounds semantics onto pixels.

Synthetic data engine

Real-world ego–exo data is still noisy and inaccurate, so we build a generative synthetic engine with exact camera poses, depth, and segmentation: 600+ interior scenes, 500+ rigged characters, 3,000 animations20K+ ego–exo training clips. The hundreds of characters are generated in a few hours.

Synthetic data engine overview
Synthetic exo rendering
Synthetic exo + ego + depth

Qualitative comparisons

If the videos don't play, please click the video boxes or change to a different browser.
Exo input
Vista4D
EgoX
Ours
GT
Cooking — IIITH 43-1
Cooking — IIITH 71-6
Cooking — IIITH 96-2
CPR — NUS 08-1

Qualitative comparison on EgoExo4D (New Activities).

If the videos don't play, please click the video boxes or change to a different browser.
Exo input
Vista4D
EgoX
Ours
GT
CPR — NUS 27-2
CPR — NUS 47-1
Basketball — SFU 05-8
COVID test — SFU 002-12

Qualitative comparison on EgoExo4D (New Activities).

If the videos don't play, please click the video boxes or change to a different browser.
Exo input
Vista4D
EgoX
Ours
GT
Omelet — UTokyo 5-1001-2
Cooking — Minnesota 040-4
Bike repair — Indiana 01-1
Cooking — Minnesota 040-2

Qualitative comparison on EgoExo4D (Unseen).

In-the-wild examples

Ego video colors are re-harmonized to the exo input for visualization purposes only.

If the videos don't play, please click the video boxes or change to a different browser.
Exo input
Ours
Laofangu
Home
BB

Chunk-autoregressive Exo2Ego (work in progress)

We also share our on-going experiments in chunk-autoregressive video generation, showing that a short finetuning stage lets the model continue from its own output and roll out much longer takes.

Quantitative results

Main results

Quantitative comparison on EgoExo4D. Best in bold, second-best underlined. Evaluated on the samples using EgoX's evaluation protocol. Unseen indicates brand new environments not seen during training. New Action (renamed from EgoX's Seen) indicate unseen actions but in environment seen during training. Grey rows show the change over EgoX. Our numbers are averaged over five independent generations. Loc Err and Contour are normalized by the output side length. Values in [brackets] are metrics the EgoX paper does not report, so we measured them ourselves.

Method Image Metrics Object Metrics Video Metrics
PSNR ↑ SSIM ↑ LPIPS ↓ CLIP-I ↑ Loc Err ↓ IoU ↑ Contour ↓ FVD ↓ T-LPIPS ↓
New Action  (new actions, seen environments)
Exo2Ego-V14.530.3840.5690.7740.3260.074622.47
TrajectoryCrafter13.050.3750.6060.7800.2100.128546.09
Wan-Fun-Control12.250.4630.6170.8100.2350.076595.07
Wan-VACE12.950.4130.6260.8290.2280.114508.69
Vista4D10.390.3010.6700.7960.1390.2540.090647.514.40
EgoX16.050.5560.4980.8960.1290.363[0.071]184.47[2.688]
Ours18.640.5600.3460.9100.0310.5760.027119.91.454
  Δ vs EgoX+2.59 dB+0.7%−30.5%+1.6%−76.0%+58.7%−62.0%−35.0%−45.9%
Unseen  (new actions, new environments)
Exo2Ego-V12.700.4390.5970.6790.4470.0031283.50
TrajectoryCrafter12.240.2970.6190.7780.4000.039821.71
Wan-Fun-Control13.590.4390.6040.7990.3990.042968.78
Wan-VACE12.170.3450.6380.8200.4000.0381045.45
Vista4D10.680.2550.6680.8110.1460.2290.087794.092.91
EgoX14.380.4570.5520.8770.3120.092[0.083]440.64[2.375]
Ours16.050.4600.4670.8940.0520.4120.044385.81.485
  Δ vs EgoX+1.67 dB+0.7%−15.4%+1.9%−83.3%+347.8%−47.0%−12.4%−37.5%

Ablation study

Ablation studies on EgoExo4D. Variants are named by what is removed relative to the full model Ours: obj tok are the per-object semantic tokens, obj mask the reprojected segmentation masks that ground them, and synth the synthetic training data. Results are averaged over five independent generations. Given ego attn mask measures Ours with attention masks taken from the ground-truth egocentric video instead of reprojected from the exo video, so it is an upper bound rather than a deployable configuration. Rows marked * are our own re-implementations of EgoX, ported to the LTX-2.3 backbone so that the comparison isolates one component at a time; EgoX's released code targets a different backbone, so these were rebuilt to the best of our ability from the paper and the reference implementation, and they may understate what the original authors would achieve. The unmarked EgoX row carries the numbers reported in their paper.

Variant Image Metrics Object Metrics Video Metrics
PSNR ↑ SSIM ↑ LPIPS ↓ CLIP-I ↑ Loc Err ↓ IoU ↑ Contour ↓ FVD ↓ T-LPIPS ↓
New Action  (new actions, seen environments)
EgoX16.050.5560.4980.8960.1290.363[0.071]184.47[2.688]
EgoX (LTX, data)*17.050.5060.4730.8750.0650.3360.061233.43.115
EgoX (LTX, data, reloc)*17.950.5320.4030.8890.0450.4560.041169.62.354
Ours (no obj tok, mask, synth)17.690.5290.4000.8940.0400.5010.036163.91.593
Ours (no obj mask, synth)18.060.5410.3770.9000.0350.5400.031142.81.540
Ours (no synth)18.390.5520.3550.9080.0320.5690.028122.91.417
Ours18.640.5600.3460.9100.0310.5760.027119.91.454
Given ego attn mask18.940.5680.3300.9130.0260.6050.022113.81.437
Unseen  (new actions, new environments)
EgoX14.380.4570.5520.8770.3120.092[0.083]440.64[2.375]
EgoX (LTX, data)*15.500.4390.5330.8530.0850.2430.078534.62.693
EgoX (LTX, data, reloc)*15.920.4490.4970.8680.0700.3160.063467.42.063
Ours (no obj tok, mask, synth)15.670.4450.4910.8820.0620.3740.054417.91.598
Ours (no obj mask, synth)15.760.4510.4810.8830.0580.3850.050399.91.596
Ours (no synth)16.000.4600.4620.8920.0520.4180.046372.41.416
Ours16.050.4600.4670.8940.0520.4120.044385.81.485
Given ego attn mask16.300.4670.4530.8970.0460.4510.038373.61.462

Acknowledgements

We thank Maksim Eisenstein, Milos Hasan, Umar Iqbal, Christian Jacobsen, Pekka Janis, Christian Laforte, Jiefeng Li, Edward Liu, and Juho Marttila for valuable discussions and help.

Citation

arXiv link coming soon — please cite the project page for now.

@online{groundedexo2ego2026,
  title  = {Grounded-Exo2Ego: Structured Semantic Grounding for
            Robust Exocentric-to-Egocentric Video Generation},
  author = {Wang, Shengze and Stengel, Michael and Li, Tianye and
            Park, Seonwook and Mazumdar, Amrita and Nagano, Koki
            and Trevithick, Alex and De~Mello, Shalini},
  year   = {2026},
  url    = {https://research.nvidia.com/labs/amri/projects/grounded-exo2ego/},
}