Scholay

学术搜索 · AI 审稿 · LaTeX 协作

DebateNav: Structured Multi-VLM Expert Debate for Robust Zero-Shot Object Navigation

作者:Henghui Sun, Weixing Tan, Lei Liu, Zhongmin Yan, Xudong Lu, Hongjun Dai · 年份:2025 · DOI:10.1109/hpcc67675.2025.00068 · 被引用次数:1 · 研究领域:Multimodal Machine Learning Applications、Reinforcement Learning in Robotics、Advanced Neural Network Applications

Zero-shot object navigation presents a highly challenging task in embodied AI, requiring an agent to interpret natural language instructions, perceive complex visual environments, and plan actions without any task-specific training. While recent approaches have introduced large language models (LLMs) as high-level planners, they often rely on static, one-shot inference and struggle with ambiguous or partially observable scenes. This paper proposes DebateNav, a novel multi-agent decision framework that integrates multiple vision-language model (VLM) experts under the supervision of a central LLM controller. Each VLM is assigned a unique expert role (e.g., object detection, risk assessment, spatial reasoning), and together they engage in structured multi-round debates when perception conflicts arise. The LLM controller performs task decomposition, memory-guided exploration, and final arbitration based on expert arguments. To enhance the perception and decision process, DebateNav incorporates a multimodal image fusion module combining RGB, depth, and segmentation inputs, as well as a map memory and trajectory tracking system that helps avoid redundant exploration and supports long-horizon planning. The system is evaluated on a subset of the HM3D dataset with approximately 5,000 tasks, achieving a$\mathbf{5 2. 3 \%}$success rate and$\mathbf{1 5. 5}$SPL under strict zero-shot conditions. Extensive ablation studies confirm the effectiveness of the expert debate mechanism, multi-mod...