Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Adaptive reverse perturbation network for audio deepfake detection

作者:Xue Ouyang, Chunhui Wang, Bin Zhao, Hao Li · 发表于:Neurocomputing · 年份:2025 · DOI:10.1016/j.neucom.2025.131466 · 被引用次数:6 · 研究领域:Speech and Audio Processing、Image and Signal Denoising Methods、Music and Audio Processing

The growing prevalence of audio deepfakes underscores the urgent need for advanced detection frameworks capable of identifying subtle synthetic artifacts. In response to this challenge, we propose an Adaptive Reverse Perturbation Network, a novel architecture that leverages partial reversal strategies on speech segments and incorporates hierarchical feature discrepancy analysis to enhance deepfake detection. Specifically, the proposed framework employs learnable reversal modules to capture phase discontinuities and spectral anomalies, and utilizes Prime-window reversal to reveal synthetic artifacts that emerge exclusively in reversed speech. Evaluations conducted on five benchmark datasets demonstrate the superior performance of the proposed method, achieving an equal error rate of 1.98 %, representing a 39.6 % improvement over previous systems, as well as a t-DCF of 0.237. Further analysis reveals an inverse correlation between language-specific weight similarity and detection accuracy. These results validate the effectiveness of the trainable differential convolution and reverse perturbation strategies in combating the evolving threat of audio deepfakes, and provide novel insights into phonological artifact patterns associated with synthetic speech.