Scholay

学术搜索 · AI 审稿 · LaTeX 协作

A Machine Learning Framework for Cognitive Impairment Screening from Speech with Multimodal Large Models

作者:Shu-Wei Chen, Ying Tan, Wenyu Hu, Yingxi Chen, Lihua Chen, Yurou He, Weihua Yu, Yang Lü · 发表于:Bioengineering · 年份:2026 · DOI:10.3390/bioengineering13010073 · 被引用次数:1 · 研究领域:Voice and Speech Disorders、Speech Recognition and Synthesis、Emotion and Mood Recognition

Background: Early diagnosis of Alzheimer’s disease (AD) is essential for slowing disease progression and mitigating cognitive decline. However, conventional diagnostic methods are often invasive, time-consuming, and costly, limiting their utility in large-scale screening. There is an urgent need for scalable, non-invasive, and accessible screening tools. Methods: We propose a novel screening framework combining a pre-trained multimodal large language model with structured MMSE speech tasks. An artificial intelligence-assisted multilingual Mini-Mental State Examination system (AAM-MMSE) was utilized to collect voice data from 1098 participants in Sichuan and Chongqing. CosyVoice2 was used to extract speaker embeddings, speech labels, and acoustic features, which were converted into statistical representations. Fourteen machine learning models were developed for subject classification into three diagnostic categories: Healthy Control (HC), Mild Cognitive Impairment (MCI), and Alzheimer’s Disease (AD). SHAP analysis was employed to assess the importance of the extracted speech features. Results: Among the evaluated models, LightGBM and Gradient Boosting classifiers exhibited the highest performance, achieving an average AUC of 0.9501 across classification tasks. SHAP-based analysis revealed that spectral complexity, energy dynamics, and temporal features were the most influential in distinguishing cognitive states, aligning with known speech impairments in early-stage AD. Conclu...