Advancing Multi-Modal Beam Prediction With Cross-Modal Feature Enhancement and Dynamic Fusion Mechanism
作者:Qihao Zhu, Yu Wang, Wenmei Li, Hao Huang, Guan Gui · 发表于:IEEE Transactions on Communications · 年份:2025 · DOI:10.1109/tcomm.2025.3548021 · 被引用次数:12 · 研究领域:Laser and Thermal Forming Techniques
In millimeter-wave and terahertz band communication systems, precise beam prediction is crucial for optimizing network performance and enhancing signal transmission efficiency. Traditional beam prediction methods have primarily relied on single-modal data, which often fails to capture the comprehensive environmental information necessary for optimal accuracy. In contrast, multi-modal data-based approaches offer a more promising solution by leveraging the strengths of diverse data sources. However, many existing fusion methods are static, inadequately accounting for variations in information content across different modalities, which can hinder the full utilization of each modality’s advantages. To address these limitations, this paper proposes an advanced multi-modal beam prediction method that integrates multipath-like data augmentation (MLDA), cross-modal feature enhancement (CMFE), and an uncertainty-aware dynamic fusion mechanism. Our approach combines image and radar data to predict beam indices, dynamically adjusting the weights of different modalities to accommodate varying information densities. The proposed method employs ResNet34 for feature extraction from the multi-modal data, followed by a cross-modal feature enhancement module that aggregates complementary information from the image and radar data. Finally, the dynamic fusion mechanism integrates the predictions from the single-modal data. Experimental results demonstrate that our method significantly improves t...