Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Reducing annotation burden in physical activity research using vision language models

作者:Abram Schönfeldt, Benjamin D. Maylor, Xiaofang Chen, Ronald Clark, Aiden Doherty · 发表于:Scientific Reports · 年份:2025 · DOI:10.1038/s41598-025-21350-6 · 被引用次数:2 · 研究领域:Physical Activity and Health、Context-Aware Activity Recognition Systems、Mobile Health and mHealth Applications

Abstract Data from wearable devices collected in free-living settings, and labelled with physical activity behaviours compatible with health research, are essential for both validating existing wearable-based measurement approaches and developing novel machine learning approaches. One common way of obtaining these labels relies on laborious human annotation of sequences of images captured by body-worn cameras. The aim of this study was to investigate whether open-source vision-language models could accurately annotate activity intensity classes in wearable camera-based validation studies, thereby reducing the annotation burden. We compared the performance of three vision language models and two discriminative models on two free-living validation studies with 161 and 111 participants, collected in Oxfordshire, United Kingdom and Sichuan, China, respectively, using the Autographer (OMG Life, defunct) wearable camera. We found that the best open-source vision-language model (VLM) and fine-tuned discriminative model (DM) achieved comparable performance when predicting sedentary behaviour from single images on unseen participants in the Oxfordshire study; median F 1 -scores: VLM = 0.89 (0.84, 0.92), DM = 0.91 (0.86, 0.95). Performance declined for light [VLM = 0.60 (0.56, 0.67), DM = 0.70 (0.63, 0.79)], and moderate-to-vigorous intensity physical activity [VLM = 0.66 (0.53, 0.85); DM = 0.72 (0.58, 0.84)]. When applied to the external Sichuan study, performance fell across all inte...