IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer
作者:Zhengyu Zou, Hao Li, Kuixuan Jiao, Liu Liu, Tingyang Xiao, X J Zhou, Fangzhou Hong, Zhizhong Su, Dingwen Zhang, Ziwei Liu · 发表于:arXiv (Cornell University) · 年份:2026 · DOI:10.48550/arxiv.2607.19228 · 研究领域:Robotics and Sensor-Based Localization、Advanced Vision and Imaging、3D Shape Modeling and Analysis
Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic reconstruction and 3D-aware vision-language methods largely rely on externally extracted 2D semantic cues or loosely coupled geometry inputs, limiting unified geometry-instance learning in long dynamic scenes. In this paper, we propose IGGT4D, a streaming instance-grounded geometry Transformer for online 4D scene understanding. IGGT4D processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. This enables long-sequence feed-forward reconstruction with geometry-instance consistency in dynamic environments. To address the lack of high-quality 4D supervision, we further construct InsScene4D-147K, a large-scale dataset spanning real/synthetic and static/dynamic scenes, with RGB images, depth, poses, and temporally consistent instance masks generated by an automated geometry-guided annotation pipeline. Experiments on 3D reconstruction, pose estimation, instance spatial tracking, and open-vocabulary segmentation demonstrate that IGGT4D outperforms e...