Investigating English Language Command Recognition using Edge-Based Hardware Accelerated Neural Processing Units
作者:S. Vijayakumar, N. Sheik Hameed, P. Nithyakumar, Praveen Chandar, N Sudhakar, B Sreela · 年份:2025 · DOI:10.1109/icidca66325.2025.11280359 · 研究领域:Speech Recognition and Synthesis、Speech and Audio Processing、Music and Audio Processing
The speech recognition is a relevant technology in the way of natural human machine-interaction; in other words, machines are capable of listening and responding to spoken orders. The present paper explains how an efficient and accurate key-word spotting model is developed in command detection. Current approaches are generally limited in generalization and high inference latency, particularly on resource limited settings. To overcome these difficulties, a powerful but lightweight model built on the base of DistilHuBERT, which is a transformer architecture designed to have the shortest possible computational complexity, is suggested. The model uses pretrained embeddings and then it undergoes fine-tuning which is designed to run on the Python platform to help in its development. The dataset used to evaluate is Google Speech Commands v2 which contains more than 105,000 audio samples of 35 predefined commands. The results of the solution are 99.2 percent accurate and the inference latency is significantly reduced to approximately 12 milliseconds. Its small form and operating system allow it to be directly integrated into smart products and speech recognition systems and speech-impaired users access solutions to be responsive and reliable in the domain.