Connectionist Speech Recognition: A Hybrid Approach
作者:Hervé A. Bourlard, Nelson H. Morgan · 发表于:Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 年份:1993 · 被引用次数:1131 · 研究领域:Speech Recognition and Synthesis、Speech and Audio Processing、Neural Networks and Applications
BACKGROUNDList of Tables 5.1 Comparison of the recognition error rates (I=insertions, S=substitutions, D=deletions) obtained with 1-state phonemic HMMs with discrete emission probabilities, Gaussian emission probabilities, and outputs of a contextual MLP. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .94 6.1 Phonetic classification rates at the frame level obtained by standard approaches."Full Gaussian" refers to the case of one Gaussian with full covariance matrix per phoneme, "MLE" refers to the case of one discrete likelihood density per phoneme estimated by counting, and "MAP" refers to the case of one discrete posterior probability density estimated by counting. . . . . . . . . . . . . . . . . . . . . . .111 6.2 Phonetic classification rates at the frame level obtained from different MLPs, compared with MLE."MLPa b-cd" stands for an MLP with blocs (width of context) of (binary) input units, hidden units and output units.The size of the output layer was kept fixed at 50 units, corresponding to the 50 phonemes to be recognized. . .113 6.3 Phonetic classification rates at the frame level obtained from contextual MLPs, compared with standard likelihoods (MLE) and a posteriori probabilities (MAP).R represents the parametrization ratio, i.e., the number of parameters divided by the number of training patterns. . . . . . . . . .116 6.4 Phonetic classification rates at the frame level on SPI-COS obtained from MLPs with linear and nonlinear outputs. . . . . . . . . ...