Multi-Object Tracking
Unified and scalable multi-object tracker using hierarchical graph neural networks that process long clips across temporal scales.
Systems for understanding dynamic 3D scenes and temporal spatial relationships.
We develop systems for understanding dynamic 3D scenes and temporal spatial relationships using multi-modal sensor data, such as cameras and LiDAR. Our research focuses on data-driven approaches that leverage temporal context to solve fundamental challenges in 4D perception: tracking objects across time, completing partial observations, and understanding scenes beyond fixed vocabularies. From graph-based multi-object tracking to feed-forward structure-from-motion, our methods enable robust spatial intelligence in real-world applications.
Detection, tracking, and segmentation • Structure-from-Motion • 4D Segmentation • Autolabeling
Unified and scalable multi-object tracker using hierarchical graph neural networks that process long clips across temporal scales.
A unified framework for zero-shot LiDAR understanding that combines text-promptable 4D segmentation with feedforward object/scene completion.
End-to-end learnable framework for efficient large-scale Structure-from-Motion using novel latent global alignment with attention mechanisms.
Fast video engine that estimates camera intrinsics, camera motion, and dense near-metric depth from unconstrained videos. Robust across diverse scenarios and camera models.