Audio-visual continuous speech recognition using a coupled hidden Markov model
作者:Xiaoxing Liu, Yibao Zhao, Xiaobo Pi, Luhong Liang, Ara Nefian · 年份:2002 · DOI:10.21437/icslp.2002-123 · 被引用次数:42 · 研究领域:Speech and Audio Processing、Advanced Data Compression Techniques、Music and Audio Processing
With the increase in the computational complexity of recent computers, audio-visual speech recognition (AVSR) became an attractive research topic that can lead to a robust solution for speech recognition in noisy environments. In the audio visual continuous speech recognition system presented in this paper, the audio and visual observation sequences are integrated using a coupled hidden Markov model (CHMM). The statistical properties of the CHMM can describe the asyncrony of the audio and visual features while preserv-ing their natural correlation over time. The experimental re-sults show that the current system tested on the XM2VTS database reduces the error rate of the audio only speech recognition system at SNR of 0db by over 55%. 1.