Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Visual-and-Language Multimodal Fusion for Sweeping Robot Navigation Based on CNN and GRU

作者:Yiping Zhang, Kolja Wilker · 发表于:Journal of Organizational and End User Computing · 年份:2024 · DOI:10.4018/joeuc.338388 · 被引用次数:5 · 研究领域:Computer Science

Effectively fusing information between the visual and language modalities remains a significant challenge. To achieve deep integration of natural language and visual information, this research introduces a multimodal fusion neural network model, which combines visual information (RGB images and depth maps) with language information (natural language navigation instructions). Firstly, the authors used faster R-CNN and ResNet50 to extract image features and attention mechanism to further extract effective information. Secondly, GRU model is used to extract language features. Finally, another GRU model is used to fuse the visual- language features, and then the history information is retained to give the next action instruction to the robot. Experimental results demonstrate that the proposed method effectively addresses the localization and decision-making challenges for robotic vacuum cleaners.