Scholay

学术搜索 · AI 审稿 · LaTeX 协作

MoTIF: An end-to-end multimodal road traffic scene understanding foundation model

作者:Zihe Wang, Haiyang Yu, Changxin Chen, Zhiyong Cui, Yufeng Bi, Yilong Ren, Zijian Wang, Delan Kong, Jing Tian, Shoutong Yuan, Zhiqiang Li · 发表于:Communications in Transportation Research · 年份:2025 · DOI:10.1016/j.commtr.2025.100227 · 被引用次数:3 · 研究领域:Traffic Prediction and Management Techniques、Autonomous Vehicle Technology and Safety、Video Surveillance and Tracking Methods

Video-based road intelligent detection constitutes a critical component in modern intelligent transportation systems, serving as a crucial role for comprehensive transportation planning and emergency traffic management. Current traffic scene perception methodologies relying on conventional deep learning architectures present inherent limitations, including heavy dependence on extensive manual annotations of specific traffic scenarios and predefined rule configurations. These approaches demonstrate constrained semantic representation capacity and limited generalizability across heterogeneous traffic scenarios. To address these challenges, this study proposes a novel end-to-end multimodal foundation model architecture that jointly generates dynamic traffic event detection outcomes and semantic-rich contextual descriptions. Through integration of low-rank adaptation (LoRA) and prompt fine-tuning as parameter-efficient fine-tuning strategies, we develop the multimodal road traffic scene understanding foundation model (MoTIF), which establishes cross-modal alignment between visual patterns and textual semantics. This framework demonstrates enhanced capability in extracting salient traffic targets and generating hierarchical scene representations, significantly improving automated detection efficiency in road video analytics. Notably, MoTIF exhibits contextual reasoning capabilities for implicit traffic event interpretation. Extensive evaluations on two real-world datasets encompas...