Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Robust 3D Target Detection Based on LiDAR and Camera Fusion

作者:Miao Jin, Bing Lu, Gang Liu, Yinglong Diao, Xiwen Chen, Gaoning Nie · 发表于:Electronics · 年份:2025 · DOI:10.3390/electronics14214186 · 被引用次数:2 · 研究领域:Advanced Neural Network Applications、Advanced Data and IoT Technologies、Visual Attention and Saliency Detection

Autonomous driving relies on multimodal sensors to acquire environmental information for supporting decision making and control. While significant progress has been made in 3D object detection regarding point cloud processing and multi-sensor fusion, existing methods still suffer from shortcomings—such as sparse point clouds of foreground targets, fusion instability caused by fluctuating sensor data quality, and inadequate modeling of cross-frame temporal consistency in video streams—which severely restrict the practical performance of perception systems. To address these issues, this paper proposes a multimodal video stream 3D object detection framework based on reliability evaluation. Specifically, it dynamically perceives the reliability of each modal feature by evaluating the Region of Interest (RoI) features of cameras and LiDARs, and adaptively adjusts their contribution ratios in the fusion process accordingly. Additionally, a target-level semantic soft matching graph is constructed within the RoI region. Combined with spatial self-attention and temporal cross-attention mechanisms, the spatio-temporal correlations between consecutive frames are fully explored to achieve feature completion and enhancement. Verification on the nuScenes dataset shows that the proposed algorithm achieves an optimal performance of 67.3% and 70.6% in terms of the two core metrics, mAP and NDS, respectively—outperforming existing mainstream 3D object detection algorithms. Ablation experiments...