DietAI24 as a framework for comprehensive nutrition estimation using multimodal large language models
作者:Runze Yan, Hanqi Luo, Jiaying Lu, Darren Liu, Hannah Posluszny, Mehak Preet Dhaliwal, Janice MacLeod, Yao Qin, Carl Yang, Terry Hartman, Xiao Hu · 发表于:Communications Medicine · 年份:2025 · DOI:10.1038/s43856-025-01159-0 · 被引用次数:13 · 研究领域:Nutrition, Genetics, and Disease、Nutritional Studies and Diet、Diet and metabolism studies
Accurate dietary assessment is essential for health research. While smartphone-based food image recognition offers a convenient alternative to traditional methods, existing computer vision approaches struggle with real-world food images and analyze only basic macronutrients, limiting their utility for comprehensive nutritional research. We developed DietAI24, a framework for automated nutrition estimation from food images that combines multimodal large language models (MLLMs) with Retrieval-Augmented Generation (RAG) technology to ground the MLLM’s visual recognition in authoritative nutrition databases rather than relying on the model’s internal knowledge. In our work, we used the Food and Nutrient Database for Dietary Studies (FNDDS) as the authoritative nutrition database. Through this approach, DietAI24 enables accurate nutrient estimation without extensive data collection or model training. DietAI24 significantly outperforms existing methods when evaluated against commercial platforms and computer vision baselines using the ASA24 and Nutrition5k datasets. Performance is measured through mean absolute error (MAE). DietAI24 achieves a 63% reduction in MAE for food weight estimation and four key nutrients and food components compared to existing methods when tested on real-world mixed dishes (p < 0.05). Notably, DietAI24 estimates 65 distinct nutrients and food components, far exceeding the basic macronutrient profiles of existing solutions. DietAI24 demonstrates that integ...