Cross-Modal Generative Semantic Communications Powered by Semantic Knowledge Base
作者:Zechuan Fang, Mengying Sun, Sen Wang, Xiaodong Xu, Haixiao Gao, Jinghong Huang, Shujun Han, P. Zhang · 发表于:IEEE Transactions on Network Science and Engineering · 年份:2026 · DOI:10.1109/tnse.2025.3650269 · 被引用次数:2 · 研究领域:Wireless Signal Modulation Classification、Generative Adversarial Networks and Image Synthesis、Advanced Wireless Communication Technologies
The rapid advancement of generative artificial intelligence (GAI) has opened new avenues for semantic communication (SemCom). In this paper, we propose CG-SemCom, a unified cross-modal generative semantic communication framework powered by shared semantic knowledge bases (SKBs). CG-SemCom leverages GAI technologies as semantic extractor at the transmitter and generative reconstructor at the receiver, enabling flexible and interpretable cross-modal transmission. We further develop VCG-SemCom, a visual transmission-oriented implementation of CG-SemCom. Specifically, the transmitter employs a vision-language large model (vLLM) to extract concise semantic information in the form of visual difference description, which is transmitted via the joint source-channel coding (JSCC). At the receiver, the diffusion model (DM)-based generative reconstructor synthesizes the target image. The knowledge retrieval mechanism tailored for shared SKBs is introduced to guide semantic extraction and ensure consistency. Additionally, a deep reinforcement learning (DRL)-driven inference agent is proposed to dynamically optimize the generation process at the receiver. To address semantic noise caused by knowledge misalignment and module mismatch, a dual-level error detection and retransmission mechanism is introduced. Moreover, we propose a novel generation similarity metric to evaluate reconstruction quality without requiring access to the original image. Extensive experiments demonstrate that the pr...