Cross-Modal Knowledge Distillation With Multi-Stage Adaptive Feature Fusion for Speech Separation
作者:Cunhang Fan, Wang Xiang, Jianhua Tao, Jiangyan Yi, Zhao Lv · 发表于:IEEE Transactions on Audio Speech and Language Processing · 年份:2025 · DOI:10.1109/taslpro.2025.3533359 · 被引用次数:17 · 研究领域:Speech and Audio Processing、Speech Recognition and Synthesis、Music and Audio Processing
Although audio-visual speech separation has achieved significant advancements, it is relatively difficult to obtain audio and visual modalities simultaneously in real scenarios, often leading to the issue of missing visual modality. Actually, during the training phase of the speech separation network, the developed audio and video datasets can be fully utilized to obtain an effectual audio-visual speech separation. However, enhancing the performance of unimodal models during testing using multimodal approaches poses a challenge. To address the problem, this paper proposes a cross-modal knowledge distillation method for speech separation, which leverages a multimodal model to enhance the unimodal model via knowledge distillation. Specifically, during the training phase, a pre-trained audio-visual network is used as the teacher and the audio-only network is used as the student. Then the teacher with the additional visual input transfers knowledge to the student to enhance the performance of the student model. During the test phase, the audio-only network only conducts speech separation. In addition, to further improve the performance of the teacher model, a multi-stage adaptive feature fusion method is proposed. The global and local perspectives are used to effectively capture deep audio-visual correlations. We have conducted extensive experiments with audio-visual datasets LRS2, LRS3, and VoxCeleb2. Experimental results demonstrate that our proposed method effectively improves...