Scholay

学术搜索 · AI 审稿 · LaTeX 协作

VTLA: Vision-Tactile-Language-Action model with preference learning for insertion manipulation

作者:Chaofan Zhang, Hao Peng, Xiaoge Cao, Xiaoshuai Hao, Shaowei Cui, Shuo Wang · 发表于:Biomimetic Intelligence and Robotics · 年份:2026 · DOI:10.1016/j.birob.2026.100333 · 被引用次数:1 · 研究领域:Tactile and Sensory Interactions、Robot Manipulation and Learning

While vision-language models have advanced significantly, their application in language-conditioned robotic manipulation is still underexplored, especially for contact-rich tasks that extend beyond visually dominant pick-and-place scenarios. To bridge this gap, we introduce Vision-Tactile-Language-Action (VTLA) model, a novel framework that enables robust policy generation in contact-intensive scenarios by effectively integrating visual and tactile inputs through cross-modal language grounding. A low-cost, multi-modal dataset has been constructed in a simulation environment, containing vision-tactile-action-instruction pairs specifically designed for the fingertip insertion task. Experimental results show that the proposed model outperforms traditional imitation learning methods ( e.g. , diffusion policy) and existing multi-modal baselines (TLA/VLA), achieving over 90% success rates on unseen peg shapes. Finally, we conduct real-world peg-in-hole experiments to demonstrate the exceptional Sim2Real performance of the proposed model. For supplementary videos and results, please visit our project website: https://sites.google.com/view/vtla .