Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Cross-Corpus Multilingual Speech Emotion Recognition: Amharic vs. Other Languages

作者:E. A. Retta, R. Sutcliffe, Jabar Mahmood, Michael Abebe Berwo, Eiad Almekhlafi, S. Khan, Shehzad Ashraf Chaudhry, Mustafa Mhamed, Junlong Feng · 发表于:Applied Sciences · 年份:2023 · DOI:10.48550/arXiv.2307.10814 · 被引用次数:14 · 研究领域:Computer Science、Engineering

In a conventional speech emotion recognition (SER) task, a classifier for a given language is trained on a pre-existing dataset for that same language. However, where training data for a language do not exist, data from other languages can be used instead. We experiment with cross-lingual and multilingual SER, working with Amharic, English, German, and Urdu. For Amharic, we use our own publicly available Amharic Speech Emotion Dataset (ASED). For English, German and Urdu, we use the existing RAVDESS, EMO-DB, and URDU datasets. We followed previous research in mapping labels for all of the datasets to just two classes: positive and negative. Thus, we can compare performance on different languages directly and combine languages for training and testing. In Experiment 1, monolingual SER trials were carried out using three classifiers, AlexNet, VGGE (a proposed variant of VGG), and ResNet50. The results, averaged for the three models, were very similar for ASED and RAVDESS, suggesting that Amharic and English SER are equally difficult. Similarly, German SER is more difficult, and Urdu SER is easier. In Experiment 2, we trained on one language and tested on another, in both directions for each of the following pairs: Amharic↔German, Amharic↔English, and Amharic↔Urdu. The results with Amharic as the target suggested that using English or German as the source gives the best result. In Experiment 3, we trained on several non-Amharic languages and then tested on Amharic. The best accu...