Scholay

学术搜索 · AI 审稿 · LaTeX 协作

A Low Computation Cost Model for Real-Time Speech Enhancement

作者:Qirui Wang, Lin Zhou, Y. F. Cao, Chenjie Zhuang, Y.M. Cheng, Yuxi Deng · 年份:2024 · DOI:10.1109/icccas62034.2024.10652686 · 被引用次数:2 · 研究领域:Speech and Audio Processing、Speech Recognition and Synthesis、Advanced Adaptive Filtering Techniques

Developing speech enhancement systems for real-time scenarios has been a challenge due to the need for low computation complexity, parallel processing, and a causal structure. In this paper, we propose a speech enhancement model that works on time-frequency domain with all operations being 1D-dimensional to reduce computation cost. Specifically, the proposed model follows a U-Net structure with several conformer blocks inserted. Our evaluation on DNS Challenge and VoiceBank + DEMAND benchmarks shows that our model performs comparably to other state-of-the-art causal systems. Most importantly, the proposed model only needs 0.70G MACs when processing 16000 samples (1 second) speech signal and achieves an RTF (Real Time Factor) of 0.012, thus indicating that the model significantly reduces the computational cost.