Multimodal Knowledge Graph Completion Using CLIP-Enhanced Hyper-Node Relational Graph Attention Networks
作者:Shashank R B, Saurabh B Swami, Akanksha A Pai, Pranamya Mady, Sharon Thomas, Minal Moharir · 发表于:International Conference Electronic Systems, Signal Processing and Computing Technologies [ICESC-] · 年份:2024 · DOI:10.1109/icesc60852.2024.10689933 · 被引用次数:1
Knowledge graph (KG) completion is a critical task in Artificial Intelligence (AI) focused on deducing absent connections between entities. The diverse nature of data, encompassing text, images, and numerical information, introduces considerable difficulties in precisely forecasting these missing links. Multi-modal KGs are necessary to fully capture the diverse and rich information associated with entities, enhancing applications in search, recommendation, and AI. This study presents a new framework for completing multi-modal knowledge graphs that integrates textual, visual, and numerical data. Pre-trained Contrastive Language-Image Pre-training (CLIP) embeddings are leveraged for textual and visual information and a dense layer is utilized for numerical data, creating a unified entity representation through a low-rank multi-modal fusion approach. The proposed model employs a Hyper-node Relational Graph Attention Network (HRGAT) to effectively aggregate and process multi-modal information. The research shows that this new approach performs significantly better than current models in certain metrics, such as Mean Rank (MR), particularly excelling in lowering the mean rank of head entities, showcasing the enhanced capability and accuracy of the method proposed in completing multi-modal KGs.