Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Multiscale Spatial Frequency-Aware Transformer and Saturation Analysis for Universal Deepfake Detection

作者:Kaiwen Xu, Xiyuan Hu, Chen Chen, Yichao Zhou, Yang Xu, Hadeel Alsolai, Shahid Mumtaz · 发表于:IEEE Transactions on Cybernetics · 年份:2026 · DOI:10.1109/tcyb.2026.3690601 · 被引用次数:1 · 研究领域:Generative Adversarial Networks and Image Synthesis、Digital Media Forensic Detection、Adversarial Robustness in Machine Learning

The rapid advancement of generative artificial intelligence technologies has introduced new security challenges, raising significant concerns. As deepfake technology becomes more sophisticated, it might be exploited by malicious actors to generate highly realistic fake images, thereby compromising the authenticity and reliability of the original content. Driven by this concern, this article introduces a universal deepfake detection model, multiscale spatial frequency-aware transformer and saturation analysis (MSFTSA), based on a multiscale, spatial-frequency-aware Transformer and saturation analysis. Unlike existing methods, MSFTSA reexamines fundamental differences between real and fake images across the frequency, spatial, and saturation domains. For this purpose, an efficient multiscale frequency-domain decoupling module has been designed to capture image features from different frequency bands, assisting in identifying inherent fake characteristics across multiple frequency scales. In addition, a spatial scattering module (SSM) is introduced to model global relationships between multiscale frequency features, achieving full-frequency interactive learning in the spatial domain. Furthermore, image saturation is used as a critical indicator to distinguish between real and fake images. Extensive experiments across multiple deepfake image datasets generated by generative adversarial networks (GANs) and diffusion models (DMs) demonstrate MSFTSA's performance, significantly outp...