Scholay

学术搜索 · AI 审稿 · LaTeX 协作

SIR-Progressive Audio-Visual TF-Gridnet with ASR-Aware Selector for Target Speaker Extraction in MISP 2023 Challenge

作者:Zhongshu Hou, Tianchi Sun, Yuxiang Hu, Changbao Zhu, Kai Chen, Jing Lü · 年份:2024 · DOI:10.1109/icasspw62465.2024.10626417 · 被引用次数:2 · 研究领域:Speech Recognition and Synthesis、Speech and Audio Processing、Music and Audio Processing

TF-GridNet has demonstrated its effectiveness in speech separation and enhancement. In this paper, we extend its capabilities for progressive audio-visual speech enhancement by introducing an attention-based audio-visual fusion module and a progressive learning strategy based on the signal-to-interference ratio (SIR). The model is integrated with a prior guided source separation (GSS) process for robust target speech extraction. A subsequent automatic speech recognition (ASR)-aware selector is employed to choose the enhancement output for better ASR performance. The proposed system achieves a final character error rate (CER) of 33.18% on the evaluation set and ranks first in the ICASSP 2024 Signal Processing Grand Challenge: Multimodal Information based Speech Processing (MISP) 2023 Challenge.