Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Text2Earth: Unlocking text-driven remote sensing image generation with a global-scale dataset and a foundation model

作者:Chenyang Liu, Keyan Chen, Rui Zhao, Zhengxia Zou, Zhenwei Shi · 发表于:IEEE Geoscience and Remote Sensing Magazine · 年份:2025 · DOI:10.1109/mgrs.2025.3560455 · 被引用次数:45 · 研究领域:Geographic Information Systems Studies、Image Retrieval and Classification Techniques、Computational Physics and Python Applications

Recently, generative foundation models (GFMs) have significantly advanced large-scale text-driven natural image generation and become a prominent research trend across various vertical domains. However, in the remote sensing field, there is still a lack of research on large-scale text-to-image (text2image) generation technology. Existing remote sensing image‒text datasets are small in scale and confined to specific geographic areas and scene types. Besides, existing text2image methods have struggled to achieve global-scale, multiresolution controllability, and unbounded image generation. To address these challenges, this article presents two key contributions: the Git-10M dataset and the Text2Earth foundation model. Git-10M is a global-scale image‒text dataset consisting of 10.5 million image‒text pairs, five times larger than the previous largest one. The dataset covers a wide range of geographic scenes and contains essential geospatial metadata, significantly surpassing existing datasets in both size and diversity. Building on Git-10M, we propose Text2Earth, a 1.3 billion-parameter GFM based on the diffusion framework to model global-scale remote sensing scenes. Text2Earth integrates a resolution guidance mechanism, enabling users to specify image resolutions. A dynamic condition adaptation (DCA) strategy is proposed for training and inference to improve image generation quality. Text2Earth not only excels in zero-shot text2image generation but also demonstrates robust gene...