Scholay

学术搜索 · AI 审稿 · LaTeX 协作

VPDNet: Virtual Point Density-Aware Network for Multimodal 3-D Object Detection

作者:Binghui Yang, Tao Tao, Jianfeng Yang, Jinsheng Xiao · 发表于:IEEE Transactions on Geoscience and Remote Sensing · 年份:2025 · DOI:10.1109/tgrs.2025.3591489 · 被引用次数:2 · 研究领域:3D Surveying and Cultural Heritage、Advanced Neural Network Applications、Robotics and Sensor-Based Localization

Lidar has become a prevalent sensor for 3D object detection in autonomous driving. However, the sparse and irregular nature of point cloud data obtained from Lidar necessitates the adoption of cross-modal object detection methods to enhance detection accuracy and stability. Nonetheless, the disparate representations of point clouds and images hinder comprehensive data fusion, leading to suboptimal performance. This paper proposes a novel multimodal density-aware 3D object detection method, VPDNet, and leverages depth completion-generated virtual point clouds to address the challenges of data fusion. To mitigate the interference caused by inaccurate depth completion, we introduce a virtual point cloud enhancement module. This module utilizes weighted features of virtual point clouds based on image segmentation information to suppress invalid information while retaining valid information. Furthermore, to further capture information from images and point clouds, we design an interactive attention fusion module that integrates features from virtual point clouds and real point clouds by adjusting attention weights. Additionally, the characteristics of Lidar dictate that collected point cloud data varies unevenly with distance, and point density, an essential feature, is often overlooked. Therefore, we propose a density-aware module that combines fused voxel features with kernel density estimation and point density information to extract spatial local features. Experiments conducte...