Application of Automatic Speech Recognition Systems in Telehealth Centers
作者:Shu-Fang Liu, Bi-Lian Chen, Yu-Li Ding, Yi‐Jui Yeh, Min-Shian Wang, Pei-Yu Su · 年份:2025 · DOI:10.1109/aixmhc65380.2025.00036 · 研究领域:Nursing Diagnosis and Documentation、COVID-19 diagnosis using AI、Speech Recognition and Synthesis
Background: The Telehealth Center at our hospital was established in August 2022. Its main task is to provide telephone consultation services to the public. According to the monthly statistics of the center's consultation volume, an average of 2,105 calls are handled each month. One of the primary impacts is on the time and quantity of nursing staff required to document nursing records.Objective: This study aims to introduce automatic speech-to-text recognition systems combined with large language model technologies to improve the completeness of nursing records and reduce transcription errors. It also seeks to enhance the usability of these systems for nursing staff.Methods: A retrospective study design was employed, analyzing 40 hours of telephone consultation data collected in 2022. The research utilized a training platform equipped with a GPU (NVIDIA A6000, 48GB). The automatic speech recognition (ASR) base model used was Whisper-medium, with a batch size of 16, a learning rate of 2e-5, and trained over 30 to 60 epochs. The loss functions applied were Binary Cross Entropy (BCE) Loss and Dice Loss. For language model prompting, the Google Gemma3-27B model was utilized as the large language model (LLM) tool. The data underwent preprocessing, including cleaning and language categorization, followed by speech recognition and model training. A total of 47 test items served as the gold standard for validity evaluation. Transcriptions were generated using Whisper’s speech-to-tex...