Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Position: Prospective of Autonomous Driving - Multimodal LLMs, World Models, Embodied Intelligence, AI Alignment, and Mamba

作者:Yunsheng Ma, Wenqian Ye, Can Cui, Haiming Zhang, Shuo Xing, Fucai Ke, Jinhong Wang, Chenglin Miao, Jintai Chen, Hamid Rezatofighi, Zhen Li, Guangtao Zheng, Chao Zheng, Tianjiao He, Manmohan Chandraker, Burhaneddin Yaman, Xin Ye, Hang Zhao, Xu Cao · 年份:2025 · DOI:10.1109/wacvw65960.2025.00114 · 被引用次数:18 · 研究领域:Artificial Intelligence in Law

With the emergence of Generative AI, multimodal AI systems that leverage foundation models are beginning to demonstrate enormous potential for perceiving the real world, collecting new data, making decisions, and using tools like humans. In recent years, the use of Large Language Models and World Models in autonomous driving has received widespread attention. However, despite their enormous potential, there is still a lack of comprehensive understanding regarding the key challenges, opportunities, and future applications of these new foundation models in driving systems. In this paper, we provide an outlook on this field, summarizing existing methods and exploring their limitations. In addition, we further discuss the applicability of emerging approaches, such as Reinforcement Learning from Human Feedback and Mamba for applications in autonomous driving. Finally, we highlight open questions and offer insights into promising directions for future research. This paper is part of a living document that will be updated based on the LLVM-AD workshop series to reflect the latest developments in the field.