Scholay

学术搜索 · AI 审稿 · LaTeX 协作

GeoPPO—A Location-Allocation Method of Superstores Based on Deep Reinforcement Learning—A Case Study of Xi’an

作者:Yuxuan Hu, Kun Qin, Shaohua Wang · 发表于:ISPRS International Journal of Geo-Information · 年份:2026 · DOI:10.3390/ijgi15030114 · 研究领域:Facility Location and Emergency Management、Vehicle Routing Optimization Methods、Urban and Freight Transport Logistics

Urban commercial restructuring, driven by the closure of traditional supermarkets and the expansion of new-format superstores, creates a large-scale spatial reallocation challenge requiring scientific location-allocation methods. Traditional heuristic algorithms such as Genetic Algorithm (GA) struggle with discrete spatial optimization under 400+ candidate sites and complex geographic mask constraints: they converge slowly and easily fall into local optima. This study proposes a Deep Reinforcement Learning (DRL) framework named GeoPPO (Geospatial Proximal Policy Optimization) to address this gap. Using Xi’an’s retail restructuring as a case setting—427 candidate locations and multidimensional geographic features—the approach models spatial constraints via a gridded environment encoded as a five-channel state tensor. Key innovations include a dynamic action-constraint mechanism that masks invalid actions based on boundary rules and competition avoidance, and a curriculum learning strategy that enables stable convergence. The framework fills the need for methods that handle hard spatial constraints in large-scale location-allocation. Tests demonstrate rapid convergence within 1,000 epochs, achieving 75% average demand coverage—2.7% and 5.5% higher than GA and Particle Swarm Optimization (PSO), respectively. Ablation experiments confirm that Vanilla PPO without dynamic action masking fails to produce feasible solutions. The framework offers a feasible technical path for handling...