Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Automatic pain face analysis in mice: Applied to a varied dataset with non-standardized conditions

作者:Niek Andresen, Manuel Wöllhaf, Jenny Wilzopolski, Annemarie Lang, Angelique Wolter, Laura Howe-Wittek, Clara Bekemeier, Leonie-Iva Pawlak, Sarah Beyer, Holger Cynis, Edith Hietel, Vera Rieckmann, Max Rieckmann, C Thöne-Reineke, Lars Lewejohann, Olaf Hellwich, Katharina Hohlbaum · 发表于:bioRxiv (Cold Spring Harbor Laboratory) · 年份:2026 · DOI:10.64898/2026.02.16.706098 · 被引用次数:1 · 研究领域:Veterinary Pharmacology and Anesthesia、Animal Behavior and Welfare Studies、Neuroendocrine regulation and behavior

Abstract Biomedical research relies on scientifically validated tools to assess pain, suffering, and distress in laboratory animals to ensure their well-being. In mice, the most frequently used laboratory animals, the Mouse Grimace Scale (MGS) provides a reliable tool for the assessment of facial expression changes caused by impaired well-being. However, no automated tool can yet reliably assess all features of the MGS across different mouse strains under varying experimental or housing conditions in real-time, as the variability present in recorded image datasets poses substantial challenges for computer vision models. Despite this technical difficulty, variability across subsets in terms of mouse strain, treatments, laboratory, and image acquisition setup is essential for paving the way toward MGS assessment under non-standardized conditions in the home cage rather than standardized cage-side recording setups. Against this background, a large and diverse dataset containing five subsets is introduced and a deep learning model was trained to predict the average MGS scores ranging between 0 and 2. It achieved a root mean squared error (RMSE) of 0.26 when trained on all subsets of the dataset, outperforming the average human rater in terms of error magnitude. The correlation between human raters and automated MGS scores was very high (Pearson’s r=0.85). In the cross-dataset evaluation, one subset was excluded from training and used for testing the model. This approach yielded h...