Listen, Perceive, Grasp: CLIP-Driven Attribute-Aware Network for Language-Conditioned Visual Segmentation and Grasping
作者:Jialong Xie, Jin Liu, Saike Huang, Chaoqun Wang, Fengyu Zhou · 发表于:IEEE Transactions on Automation Science and Engineering · 年份:2024 · DOI:10.1109/tase.2024.3510777 · 被引用次数:2 · 研究领域:Multimodal Machine Learning Applications、Visual Attention and Saliency Detection、Advanced Neural Network Applications
Endowing robots with the ability to understand natural language and execute grasping is a challenging task in a human-centric environment. Existing works on language-conditioned grasping achieve end-to-end grasping detection based on language. However, these works lack fine-grained visual grounding, resulting in cognitive deficits for robots. Moreover, they ignore the correlation between visual attributes of objects and grasping, leading to coarse grasp poses. To this end, we propose a CLIP-driven aTtribute-aware network (CTNet) for language-conditioned visual segmentation and grasping, enabling the robots to listen, perceive, and grasp the referred object in real-world applications. Specifically, we first employ Listen stage to understand basic linguistic and visual concepts. Subsequently, we introduce Perceive stage to mine multi-modal features and visual attribute cues (e.g., boundary and spatial location), then yield a language-conditioned segmentation mask. Further, we design Grasp stage to aggregate the perceived attribute information and refine the spatial location and grasping rectangle, generating a high-quality grasp pose. Lastly, we provide an extended large dataset Ref-OCID-Grasp to train and test our method, achieving a grasping accuracy of 97.76% and segmentation OIoU of 91.82%. The real-world robotic applications demonstrate the effectiveness of our proposed approach. The project, video, and dataset can be found athttps://ctnetgrasp.github.io. Note to Practitio...