Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
作者:Sanyuan Chen, Chengyi Wang, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Tie‐Yan Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, Furu Wei · 发表于:IEEE Transactions on Audio Speech and Language Processing · 年份:2025 · DOI:10.1109/taslpro.2025.3530270 · 被引用次数:112 · 研究领域:Speech Recognition and Synthesis、Natural Language Processing Techniques、Speech and dialogue systems
We introduce a language modeling approach for text to speech synthesis (TTS). Specifically, we train aneural codec language model(calledVALL-E) using discrete codes derived from an off-the-shelf neural audio codec model, and regard TTS as a conditional language modeling task rather than continuous signal regression as in previous work. During the pre-training stage, we scale up the TTS training data to 50 k hours of English speech which is hundreds of times larger than existing systems.VALL-Eemergesin-context learningcapability and can be used to synthesize high-quality personalized speech with only a 3-second enrolled recording of an unseen speaker as a prompt. Experiment results show thatVALL-Esignificantly outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity. In addition, we findVALL-Ecould preserve the speaker's emotion and acoustic environment from the prompt in synthesis.