Scholay

学术搜索 · AI 审稿 · LaTeX 协作

GFMLLM: Enhance multi-modal large language model for global and fine-grained visual spatial perception

作者:Zhendong Fan, Cheng Zhang, Jianqiong Gao, Hao Wei, Hongbo Gao, Tao Xie, Ruifeng Li, Li-Jun Zhao · 发表于:Expert Systems with Applications · 年份:2025 · DOI:10.1016/j.eswa.2025.130239 · 研究领域:Multimodal Machine Learning Applications、Speech and dialogue systems、Topic Modeling