Learning and Perception Research

GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images

GATOR reconstructs a strainer's open mesh and detailed handle, and composes tabletop and indoor objects using their predicted scene-relative poses.
01 / Overview

Three Ways to Reconstruct the Same Object

Generation + refinement (GATOR), generation only (Pixal3D), and an agent working from scratch (GPT-6), using the same input images. A qualitative water-pouring simulation compares GATOR and Pixal3D.
02 / Method

Generate. Inspect. Refine.

Correspondence-aligned observations guide generation. Evidence-driven edits refine the resulting asset.

GATOR pipeline: correspondence-aligned modality mixing, structure flow, geometry flow, appearance flow, and an evidence-driven agentic refinement loop.
RGB, target masks and pointmaps jointly condition sparse structure generation. Category text guides sparse structure; projected image features and object descriptions guide geometry and appearance. The generated asset then initializes local, evidence-driven refinement.
01

Pose-aware modality mixing

Local attention mixes patch-aligned RGB, target-mask and pointmap features before cross-view reasoning, enabling complete object reconstruction in clutter with scene-relative pose.

02

Text-guided generation

Stage-specific adapters use category names for structure and object descriptions for geometry and appearance, complementing projected image features to recover detailed textures and physically based materials.

03

Agentic asset refinement

An image-grounded edit-render-review loop repairs structure, topology and textures while preserving reliable geometry and pose, improving simulation readiness without retraining the generator.

03 / Agentic refinement

Local Edits. Visible Evidence.

Given the original images, masks and available geometry, an agent edits the generated mesh in Blender. It restores fine structure, repairs topology, completes missing parts and corrects textures, while preserving reliable regions and the asset's coordinate frame.

Original and candidate assets are rendered under matched cameras and lighting. The agent compares them, retains supported edits and reverts unsuccessful ones. Each object has a ten-minute work budget with no fixed revision count; mesh export is timed separately.

Six input, before and after examples: recovering a bus railing, correcting package text, completing chair legs, removing extraneous desk geometry, repairing an umbrella handle and preserving an already accurate table.
Paired renders expose local changes in geometry and appearance. Refinement can also leave an already accurate reconstruction largely unchanged.

The agent has no access to reference meshes, held-out observations or benchmark scores. All delivered assets are evaluated without quality filtering, including unchanged outputs.

04 / Interactive reconstructions

Explore the reconstructions in 3D

22 objects from Figures 3, 4, 9 and 10. GATOR includes agentic refinement; GPT-6 reconstructs from scratch. Assets are centered and scaled for inspection in the figures' display orientations. Interactive views use consistent diffuse shading to compare shape and texture.

Loading 3D viewers

Composed scene reconstructions

Reconstructed objects retain their predicted positions, orientations and relative scales. Explore three ScanNet++ rooms and four HANDAL scenes with 63 refined objects and their input views. Compare against ShapeR, SimFoundry, RecGen and MV-SAM3D with linked cameras and a shared scene scale. Toggle the estimated point cloud to inspect placement. ShapeR provides geometry only; SimFoundry uses one view and RecGen uses up to two. Available object counts are shown for each method. Green outlines mark the reconstructed objects in each image.

Loading scene viewer
05 / Analysis

A Strong Starting Point Matters.

Quality under limited agent time

On ten sampled Toys4K objects, GATOR without refinement takes about 30 seconds and already outperforms the from-scratch agent with a twenty-minute budget on chamfer distance and LPIPS. The first point is generation only. The remaining refinement points use mean recorded agent time plus an estimated 32 seconds for generation, excluding mesh export.

Quality versus time curves beginning with GATOR generation only at 0.5 minutes, alongside an example and results at two-, five-, ten- and twenty-minute agent budgets.
Means include valid outputs only: the from-scratch agent returns 0/10 valid assets at one minute, 5/10 at two minutes and 10/10 from five minutes onward. GATOR returns valid assets at every tested budget. The first example column shows our generator-only result at approximately 0.5 minutes; the other headings denote agent budgets.

From reconstruction to rigid-body simulation

Reconstructed assets form complete scenes for physics simulation and robot interaction. The videos show a conference room settling under gravity and three humanoid robots interacting with the reconstructed environment.

Conference-room settling. Input views and a 46-object reconstructed scene settling in Isaac Sim.
Humanoid interaction. Humanoids bump into clutter, lift and reposition a chair from a stack, and pick up a waste basket to relocate it to a room corner.