Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives

作者:Yang Zhou, Junjie Li, Chun‐Quan Ou, Dawei Yan, Haokui Zhang, Xizhe Xue · 发表于:Drones · 年份:2025 · DOI:10.3390/drones9080557 · 被引用次数:10 · 研究领域:Advanced Neural Network Applications、Advanced Image and Video Retrieval Techniques、Multimodal Machine Learning Applications

Due to its extensive applications, aerial image object detection has long been a hot topic in computer vision. In recent years, advancements in unmanned aerial vehicle (UAV) technology have further propelled this field to new heights, giving rise to a broader range of application requirements. However, traditional UAV aerial object detection methods primarily focus on detecting predefined categories, which significantly limits their applicability. The advent of cross-modal text–image alignment (e.g., CLIP) has overcome this limitation, enabling open-vocabulary object detection (OVOD), which can identify previously unseen objects through natural language descriptions. This breakthrough significantly enhances the intelligence and autonomy of UAVs in aerial scene understanding. This paper presents a comprehensive survey of OVOD in the context of UAV aerial scenes. We begin by aligning the core principles of OVOD with the unique characteristics of UAV vision, setting the stage for a specialized discussion. Building on this foundation, we construct a systematic taxonomy that categorizes existing OVOD methods for aerial imagery and provides a comprehensive overview of the relevant datasets. This structured review enables us to critically dissect the key challenges and open problems at the intersection of these fields. Finally, based on this analysis, we outline promising future research directions and application prospects. This survey aims to provide a clear road map and a valuabl...