Animal acoustic identification, denoising and source separation using generative adversarial networks
作者:Mei Wang, Kevin Darras, Renjie Xue, Fanglin Liu · 发表于:Methods in Ecology and Evolution · 年份:2025 · DOI:10.1111/2041-210x.70148 · 被引用次数:7 · 研究领域:Animal Vocal Communication and Behavior、Speech and Audio Processing、Music and Audio Processing
Abstract Soundscapes contain rich ecological information, offering insights into both biodiversity and ecosystem dynamics. However, the sheer volume of data produced by passive acoustic monitoring presents significant challenges for scalable analysis and ecological interpretation. While convolutional neural networks (CNNs) have advanced species classification in bioacoustics, they often struggle with identifying acoustic targets in acoustic space and quantifying soundscapes' characteristics. In this study, we propose a novel spectrogram‐to‐spectrogram translation framework based on generative adversarial networks (GANs) to isolate and quantify acoustic sources within soundscape recordings. Our method is trained on paired spectrogram images: original full‐spectrogram representations and target spectrogram representations containing only the vocalizations of specific sound labels. This design enables the model to learn source‐specific mappings and perform both the species and community‐level separation of acoustic components in soundscape recordings. We developed and evaluated two GAN‐based models: a species‐level GAN targeting eight avian species, and a community‐level GAN distinguishing among avian, insect and anthropogenic sound sources. The models were trained and tested using soundscape recordings collected from the Yaoluoping National Nature Reserve, eastern China. The species‐level model achieved a mean F1 score of 0.76 for pixel‐wise detection, while the community‐level...