Scholay

学术搜索 · AI 审稿 · LaTeX 协作

A deepfake detection framework based on multimodal dynamic fusion and semantic consistency analysis

作者:Aodi Chen, Qianru Lin · 发表于:International Conference on Telecommunications, Optics and Computer Science · 年份:2026 · DOI:10.1117/12.3108700 · 研究领域:Engineering

Existing Deepfake detection studies predominantly focus on single-modality artifacts, yet such approaches exhibit limited generalization when facing fast-evolving generative models and often fail against high-quality forgeries. More critically, they overlook semantic inconsistencies across modalities—an essential cue humans naturally rely on for deception detection and a primary channel through which AI-generated misinformation exerts real-world influence. To bridge this gap, this work proposes DF-SemFuse, a multimodal Deepfake detection framework built upon dynamic cross-modal fusion and semantic consistency analysis across video, audio, and text modalities. The central research question this paper addressed is how to effectively model and quantify the temporal interactions between multiple modalities and further distill semantic contradictions that violate real-world logic. DF-SemFuse consists of three key modules: (1) powerful single-modality encoders capturing facial motion and lip dynamics (visual), acoustic and linguistic cues (audio), and contextual semantics (text); (2) a temporal cross-modal attention and gated fusion module that adaptively assigns modality- and time-dependent importance rather than performing naive feature concatenation; and (3) a semantic consistency analysis module leveraging knowledge-graph embeddings and commonsense reasoning to detect high-level logical contradictions. Extensive experiments on multiple public datasets demonstrate that DF-SemFus...