Scholay

学术搜索 · AI 审稿 · LaTeX 协作

DFNet: A Dual LiDAR–Camera Fusion 3-D Object Detection Network Under Feature Degradation Condition

作者:Tao Ye, Ruohan Liu, Chengzu Min, Yuliang Li, Xiaosong Li · 发表于:IEEE Sensors Journal · 年份:2025 · DOI:10.1109/jsen.2025.3551149 · 被引用次数:4 · 研究领域:Advanced Neural Network Applications、Industrial Vision Systems and Defect Detection、Advanced Image and Video Retrieval Techniques

LiDAR-camera fusion is widely used in 3-D perception tasks. In LiDAR and camera sensing tasks, the hierarchical feature abstraction capability possessed by the deep network is beneficial to capture the detailed information from point clouds and RGB images. However, it tends to filter some of the information to extract important features, where the problem of feature degradation due to loss of useful information is inevitable. The deterioration of LiDAR-camera fusion due to feature degradation, brought about by this factor, becomes a challenging problem. It reduces object recognition and leads to decreased detection accuracy. To address this problem, we propose a dual LiDAR-camera fusion network (DFNet) based on cross-modal compensation and feature enhancement. We design a multimodal feature extraction (MFE) module to complement the sparse features of the point cloud utilizing image features and focusing on the spatial information of the features. Then, we introduce a multiscale feature aggregation (MFA) module to generate bird’s-eye view (BEV) representations of the features, which generates feature proposals that are then input to the voxel-grid aggregation (VGA) module to obtain the grid-pooled features. Meanwhile, the VGA module receives the feature proposals extracted from the image backbone and projects the point cloud through voxels to obtain voxel-fused features. Finally, we aggregate the grid-pooled features and voxel-fused features to produce more informative fused f...