Toward Visual Grounding: A Survey
作者:Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan, Yaowei Wang, Changsheng Xu · 发表于:IEEE Transactions on Pattern Analysis and Machine Intelligence · 年份:2025 · DOI:10.1109/tpami.2025.3630635 · 被引用次数:9 · 研究领域:Multimodal Machine Learning Applications、Subtitles and Audiovisual Media、Visual Attention and Saliency Detection
Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) based on the given expression text. This task simulates the common referential relationships between visual and linguistic modalities, enabling machines to develop human-like multimodal comprehension capabilities. Consequently, it has extensive applications in various domains. However, since 2021, visual grounding has witnessed significant advancements, with emerging new concepts such as grounded pre-training, grounding multimodal LLMs, generalized visual grounding, and giga-pixel grounding, which have brought numerous new challenges. In this survey, we first examine the developmental history of visual grounding and provide an overview of essential background knowledge, including fundamental concepts and evaluation metrics. We systematically track and summarize the advancements, and then meticulously define and organize the various settings to standardize future research and ensure a fair comparison. In the dataset section, we compile a comprehensive list of current relevant datasets, conduct a fair comparative analysis, and provide ultimate performance prediction to inspire the development of new standard benchmarks. Additionally, we delve into numerous applications and highlight several advanced topics. Finally, we outline the challenges confronting visual grounding and propose valuable directions for future research, which may s...