SpatialScene2Vec: A self-supervised contrastive representation learning method for spatial scene similarity evaluation
作者:Danhuai Guo, Yingxue Yu, Shiyin Ge, Song Gao, Gengchen Mai, Huixuan Chen · 发表于:International Journal of Applied Earth Observation and Geoinformation · 年份:2024 · DOI:10.1016/j.jag.2024.103743 · 被引用次数:17 · 研究领域:Advanced Image and Video Retrieval Techniques、Multimodal Machine Learning Applications、Geographic Information Systems Studies
Spatial scene similarity plays a crucial role in spatial cognition, as it enables us to understand and compare different spatial scenes and their relationships. However, understanding spatial scenes is a complex task. While existing literature has contributed to spatial scene representation learning, these methods primarily focus on comprehending the spatial relationships among objects, often neglecting their semantic features. Furthermore, there is a lack of scene representation learning methods that can seamlessly handle different types of spatial objects (e.g., points, polylines, and polygons) in a scene. Moreover, since expert knowledge is required for the annotation process of spatial scene understanding, publicly available high-quality annotation data has a limited size which usually leads to suboptimal results. To address these issues, we propose a novel multi-scale spatial scene encoding model called SpatialScene2Vec. SpatialScene2Vec utilizes a point location encoder to seamlessly encode the spatial information of different types of spatial objects. A point feature encoder is employed to encode the semantic features of these objects. A spatial scene embedding is generated by integrating the spatial embeddings and feature embeddings of spatial objects within this scene. Furthermore, to address the limited labeled data problem, we propose a self-supervised learning framework to train the SpatialScene2Vec model in which a contrastive loss is used for spatial scene simil...