Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Unambiguous Pyramid Cost Volumes Fusion for Stereo Matching

作者:Qibo Chen, Baozhen Ge, Jianing Quan · 发表于:IEEE Transactions on Circuits and Systems for Video Technology · 年份:2023 · DOI:10.1109/tcsvt.2023.3291726 · 被引用次数:19 · 研究领域:Advanced Vision and Imaging、Advanced Image and Video Retrieval Techniques、Advanced Image Processing Techniques

Stereo matching is a challenging task in 3D vision. Only relying on single-scale cost aggregation provides deficient matching information. Prior works thus try to adopt pyramid cost volumes fusion to calculate the matching cost. However, the commonly used cost volume fusion process can not fully exploit the benefits of these multi-scale cost volumes. Motivated by the cross-scale feature discrepancy, we propose an Unambiguous Pyramid cost volumes Fusion Network terms as UPFNet, to reduce the ambiguity between pyramid cost volumes at different scales and boost the cross-scale information flow in the stereo matching framework based on 3D convolution. First, we propose a pyramid-cost progressive fusion (PPF) module, which adds consistent supervision for pre-fusion cost volumes to reduce feature semantic inconsistency and facilitates cross-scale interactions to narrow the detailed gap between different scales. The output disparity can be gradually refined in a coarse-to-fine manner. Furthermore, we design a residual disparity aggregation (RDA) module, introducing disparity dimension information to further exploit the local aggregation capability of 3D convolution by squeezing disparity and exciting channel response. Extensive experiments on the Scene Flow, KITTI and Middlebury benchmarks demonstrate the effectiveness of the proposed UPFNet. The results show that the proposed approach achieves state-of-the-art performance and is ranked first in the KITTI 2015 leaderboard when submi...