Scholay

学术搜索 · AI 审稿 · LaTeX 协作

R M 2 Occ: Re-Projection Multi-Task Multi-Sensor Fusion for Autonomous Driving 3D Object Detection and Occupancy Perception

作者:Yilong Ren, Lening Wang, Minda Li, Han Jiang, Zhiyong Cui, Mengmeng Yang, Haiyang Yu, Diange Yang · 发表于:IEEE Transactions on Intelligent Transportation Systems · 年份:2025 · DOI:10.1109/tits.2025.3606554 · 被引用次数:15 · 研究领域:Advanced Neural Network Applications、Robotics and Sensor-Based Localization、Industrial Vision Systems and Defect Detection

Occupancy prediction plays a crucial role in supporting autonomous driving planning and decision-making. Existing methods typically rely on modular stacking and fusion techniques of object detection, semantic segmentation, and depth estimation to achieve 3D occupancy. However, they fail to deeply explore the transformation relationships between 2D and 3D spaces and to efficiently fuse the different characteristics of multi-source sensors. We propose R$M^{2}$Occ, the first 3D occupancy perception network that integrates multi-sensor fusion based on different sensor principles and achieves multi-task learning. To leverage the rich 2D semantic information captured by cameras and elevate it to the 3D domain, we begin by querying and populating predefined empty voxels with multi-view image features. Subsequently, we progressively fuse 3D LiDAR point clouds with these populated voxels through an unbalanced fusion strategy that effectively supplements missing information and suppresses noise. Leveraging IMU data and calibration parameters, we then re-project the enriched voxels back onto the 2D image plane according to camera coordinates, performing a secondary query using the semantic segmentation results to recover semantic details potentially lost due to radar fusion limitations and incomplete voxel querying. Finally, supported by a multi-task detection head, R$M^{2}$Occ simultaneously accomplishes 3D object detection, semantic segmentation, Bird’s Eye View (BEV) detection, and f...