Scholay

学术搜索 · AI 审稿 · LaTeX 协作

A comprehensive experimental study for analyzing the effects of data augmentation techniques on voice classification

作者:Halit Bakır, A. Çayır, T. S. Navruz · 发表于:Multimedia tools and applications · 年份:2023 · DOI:10.1007/s11042-023-16200-4 · 被引用次数:29 · 研究领域:Computer Science

It is not always possible to find enough data for deep learning studies. So, various data augmentation techniques have been developed and thus the success of deep learning models has increased. In this work, a comprehensive study has been conducted to evaluate the efficiency of different data augmentation techniques in terms of improving the voice classification models’ performance. To this end, we proposed extracting MFCC features from the audio files, converting them into RGB images, and using CNN deep learning model for classifying the constructed RGB images into 12 classes. Moreover, Random search algorithm has been adopted for tuning the hyperparameters and selecting the best CNN model that can achieve this task with as high performance as possible. After that, some voice augmentation and image augmentation techniques have been used to increase the number of samples in the original dataset, and 17 different datasets have been constructed and used for training the proposed model. Particularly, 5 different voice augmentation techniques have been used for constructing 5 different datasets from the original dataset. When the proposed model has been trained using these voice augmentation-based datasets the validation accuracy and F1-score exceed 95%. Then, we suggested using image augmentation techniques for constructing another dataset from the original dataset. By training the model using this dataset we noted that the validation accuracy and F1-score of the proposed model ...