Conformer-based Tibetan Speech Recognition Algorithm
作者:Yongbin Yu, Qing Huang, Manping Fan, Yutong Liu, Jiaheng Ding, Xiangxiang Wang, Ziyue Zhang, Lei Li, Nuo Qun, Duojie Renzeng, Tashi Karma, Favour Ekong · 年份:2024 · DOI:10.1145/3708657.3708775 · 研究领域:Speech Recognition and Synthesis、Speech and Audio Processing、Infant Health and Development
This paper highlights a Tibetan speech recognition algorithm based on the Conformer model, aiming to improve communication efficiency in Tibetan-Chinese bilingual communities and promote the preservation and transmission of the Tibetan language. By combining convolutional neural networks and the self-attention mechanism, this algorithm effectively captures both local and global features of speech signals. Experiments conducted on a Tibetan dialect speech synthesis dataset show that the model's Word Error Rate (WER) can be reduced to 12.69%, indicating strong performance in Tibetan speech recognition tasks. The results provide new ideas for the development of Tibetan speech recognition technology and offer a pathway for future expansion of dialect data and model optimization.