Pose-aware modality mixing
Local attention mixes patch-aligned RGB, target-mask and pointmap features before cross-view reasoning, enabling complete object reconstruction in clutter with scene-relative pose.
Correspondence-aligned observations guide generation. Evidence-driven edits refine the resulting asset.
Local attention mixes patch-aligned RGB, target-mask and pointmap features before cross-view reasoning, enabling complete object reconstruction in clutter with scene-relative pose.
Stage-specific adapters use category names for structure and object descriptions for geometry and appearance, complementing projected image features to recover detailed textures and physically based materials.
An image-grounded edit-render-review loop repairs structure, topology and textures while preserving reliable geometry and pose, improving simulation readiness without retraining the generator.
Given the original images, masks and available geometry, an agent edits the generated mesh in Blender. It restores fine structure, repairs topology, completes missing parts and corrects textures, while preserving reliable regions and the asset's coordinate frame.
Original and candidate assets are rendered under matched cameras and lighting. The agent compares them, retains supported edits and reverts unsuccessful ones. Each object has a ten-minute work budget with no fixed revision count; mesh export is timed separately.

The agent has no access to reference meshes, held-out observations or benchmark scores. All delivered assets are evaluated without quality filtering, including unchanged outputs.
22 objects from Figures 3, 4, 9 and 10. GATOR includes agentic refinement; GPT-6 reconstructs from scratch. Assets are centered and scaled for inspection in the figures' display orientations. Interactive views use consistent diffuse shading to compare shape and texture.
Reconstructed objects retain their predicted positions, orientations and relative scales. Explore three ScanNet++ rooms and four HANDAL scenes with 63 refined objects and their input views. Compare against ShapeR, SimFoundry, RecGen and MV-SAM3D with linked cameras and a shared scene scale. Toggle the estimated point cloud to inspect placement. ShapeR provides geometry only; SimFoundry uses one view and RecGen uses up to two. Available object counts are shown for each method. Green outlines mark the reconstructed objects in each image.
On ten sampled Toys4K objects, GATOR without refinement takes about 30 seconds and already outperforms the from-scratch agent with a twenty-minute budget on chamfer distance and LPIPS. The first point is generation only. The remaining refinement points use mean recorded agent time plus an estimated 32 seconds for generation, excluding mesh export.

Reconstructed assets form complete scenes for physics simulation and robot interaction. The videos show a conference room settling under gravity and three humanoid robots interacting with the reconstructed environment.