DGGR-Net: single-image 3D reconstruction from complex backgrounds via graph-based refinement and difference-guided fusion
作者:Yang Ding, Huamin Yang, Chao Xu, Chao Zhang, Linxuan Li · 发表于:Journal of King Saud University - Computer and Information Sciences · 年份:2025 · DOI:10.1007/s44443-025-00251-8 · 被引用次数:3 · 研究领域:3D Shape Modeling and Analysis、Computer Graphics and Visualization Techniques、Advanced Vision and Imaging
Accurately recovering 3D shapes from single images that contain complex backgrounds remains a longstanding and difficult task. Unlike machines, the human visual system can effortlessly filter out background distractions and utilize extensive geometric and semantic knowledge to interpret 3D structures precisely. In contrast, current single-image 3D reconstruction methods often struggle to focus on the target object when confronted with complex backgrounds, as noise and irrelevant objects reduce reconstruction accuracy. To address this issue, we propose a cross-modal fusion strategy that integrates feature enhancement with difference-guided attention to enable high-quality 3D reconstruction from single complex images (named DGGR-Net). Specifically, we utilize retrieved 3D models from the ShapeNet dataset as structural priors and introduce a local geometry-preserving graph convolution module (LGPConv) to optimize fine-grained point cloud structures. Additionally, we design a bidirectional spatial attention (BSA) module to effectively capture spatial image features, reducing background interference during feature extraction. Furthermore, we propose a difference-guided cross-modal attention (DCA) module, which explicitly computes the differences between image and point cloud features to guide precise cross-modal feature fusion, thereby improving modality complementarity and robustness. The experimental results show that our proposed method achieves a Chamfer Distance (CD) of $$3.1...