Scholay

学术搜索 · AI 审稿 · LaTeX 协作

SRC-IT2: Speech Rate-Controllable Mongolian Emotional Speech Synthesis Based on Improved Tacotron2

作者:Ren Qing-dao-er-ji, Qian Bo, Chao Zhou, Yatu Ji, Nier Wu · 发表于:Electronics · 年份:2025 · DOI:10.3390/electronics14193835 · 被引用次数:4 · 研究领域:Speech Recognition and Synthesis、Speech and Audio Processing、Music and Audio Processing

To address the challenges of slow synthesis speed, unstable quality, limited emotional expressiveness, and the lack of controllable speaking rate in Mongolian emotional speech synthesis, this paper proposes a speech Rate-Controllable Mongolian emotional speech synthesis model based on improved Tacotron2 (SRC-IT2). First, an end-to-end Mongolian speech synthesis module is constructed based on an improved Tacotron2 framework, incorporating the unique linguistic characteristics of the Mongolian script. The front-end processing is optimized accordingly, and a G2P-Seq2Seq model is employed to achieve accurate grapheme-to-phoneme conversion for Mongolian characters. Next, on top of the end-to-end synthesis framework, a joint text-audio emotion analysis module is integrated to effectively learn and represent emotional style features specific to Mongolian speech. Finally, a style encoder and speaking rate control variable are embedded into the acoustic modeling process, further enhancing Tacotron2’s ability to dynamically adjust the speaking rate during emotional speech generation. Experimental results demonstrate that the proposed model produces more natural-sounding speech with improved emotional expressiveness and enables effective real-time control over speaking rate in Mongolian emotional speech synthesis.