Scholay

学术搜索 · AI 审稿 · LaTeX 协作

RSVQ-Diffusion Model for Text-to-Remote-Sensing Image Generation

作者:Xin Gao, Yao Fu, Xiaonan Jiang, Fanlu Wu, Yu Zhang, Tianjiao Fu, Chao Li, Junyan Pei · 发表于:Applied Sciences · 年份:2025 · DOI:10.3390/app15031121 · 被引用次数:6 · 研究领域:Image Retrieval and Classification Techniques、Advanced Image and Video Retrieval Techniques、Remote-Sensing Image Classification

Despite significant challenges, the text-guided remote sensing image generation method shows great potential in many practical applications such as generative adversarial networks in remote sensing tasks; generated images still face challenges such as low realism, face challenges, and unclear details. Moreover, the inherent spatial complexity of remote sensing images and the limited scale of publicly available datasets make it particularly challenging to generate high-quality remote sensing images from text descriptions. To address these challenges, this paper proposes the RSVQ-Diffusion model for remote sensing image generation, achieving high-quality text-to-remote-sensing image generation applicable for target detection, simulation, and other fields. Specifically, this paper designs a spatial position encoding mechanism to integrate the spatial information of remote sensing images during model training. Additionally, the Transformer module is improved by incorporating a short-sequence local perception mechanism into the diffusion image decoder, addressing issues of unclear details and regional distortions in generated remote sensing images. Compared with the VQ-Diffusion model, our proposed model achieves significant improvements in the Fréchet Inception Distance (FID), the Inception Score (IS), and the text–image alignment (Contrastive Language-Image Pre-training, CLIP) scores. The FID score successfully decreased from 96.68 to 90.36; the CLIP score increased from 26.92 t...