Scholay

学术搜索 · AI 审稿 · LaTeX 协作

P2-Net: Joint Description and Detection of Local Features for Pixel and Point Matching

作者:Bing Wang, Changhao Chen, Zhaopeng Cui, Jie Qin, Chris Xiaoxuan Lu, Zhengdi Yu, Peijun Zhao, Zhen Dong, Fan Zhu, Niki Trigoni, Andrew Markham · 发表于:2021 IEEE/CVF International Conference on Computer Vision (ICCV) · 年份:2021 · DOI:10.1109/iccv48922.2021.01570 · 被引用次数:64 · 研究领域:Robotics and Sensor-Based Localization、3D Surveying and Cultural Heritage、Advanced Image and Video Retrieval Techniques

Accurately describing and detecting 2D and 3D key-points is crucial to establishing correspondences across images and point clouds. Despite a plethora of learning-based 2D or 3D local feature descriptors and detectors having been proposed, the derivation of a shared descriptor and joint keypoint detector that directly matches pixels and points remains under-explored by the community. This work takes the initiative to establish fine-grained correspondences between 2D images and 3D point clouds. In order to directly match pixels and points, a dual fully-convolutional framework is presented that maps 2D and 3D inputs into a shared latent representation space to simultaneously describe and detect keypoints. Furthermore, an ultra-wide reception mechanism and a novel loss function are designed to mitigate the intrinsic information variations between pixel and point local regions. Extensive experimental results demonstrate that our framework shows competitive performance in fine-grained matching between images and point clouds and achieves state-of-the-art results for the task of indoor visual localization. Our source code is available at https://github.com/BingCS/P2-Net.