Feature Pyramid Multi-View Stereo Network Based on Self-Attention Mechanism
作者:Junjie Li, Zhengyao Bai, Wei Cheng, Huijie Liu · 年份:2022 · DOI:10.1145/3512388.3512422 · 被引用次数:3 · 研究领域:Advanced Vision and Imaging、Optical measurement and interference techniques、Image Processing Techniques and Applications
Multi-view 3D reconstruction aims to recover the 3D geometry of an object from a set of 2D images. The traditional methods match a pair of the corresponding pixels by calculating the variance, which is primarily a pixel-based measure without consideration of the interdependence of pixels and is ineffective for matching texture-free or occluded regions. To address this issue, we investigate multi-view 3D reconstruction using deep learning and propose a feature pyramid multi-view stereo network based on self-attention mechanism (FPSA-MVSNet). We use a multi-scale feature pyramid network to extract deep image features and introduce the self-attention mechanism into the hierarchical feature extraction module to capture an extensive range of interdependencies between pixels. Meanwhile, we use the group correlation method and a cascade approach to construct lightweight cost volume, and employ a coarse-to-fine depth map inference strategy to predict the initial depth map at the coarsest level, then iteratively optimize it by progressive up-sampling to generate the high-resolution depth map at the finest level. We performed experiments on the DTU benchmark dataset and compared our method with the existing multi-view 3D reconstruction methods. Experimental results show that the proposed model outperforms the existing best network model CasMVSNet on the DTU benchmark dataset, with significant improvements in completeness and overall score, by 26.5% and 9.0%, respectively.