Scholay

学术搜索 · AI 审稿 · LaTeX 协作

EDEN: Enhanced Diffusion for High-quality Large-motion Video Frame Interpolation

作者:Zihao Zhang, Haoran Chen, Haoyu Zhao, Guansong Lu, Yanwei Fu, Songcen Xu, Zuxuan Wu · 年份:2025 · DOI:10.1109/cvpr52734.2025.00202 · 被引用次数:3 · 研究领域:Advanced Vision and Imaging、Advanced Image Processing Techniques、Image Processing Techniques and Applications

Handling complex or nonlinear motion patterns has long posed challenges for video frame interpolation. Although recent advances in diffusion-based methods offer improvements over traditional optical flow-based approaches, they still struggle to generate sharp, temporally consistent frames in scenarios with large motion. To address this limitation, we introduce EDEN, an Enhanced Diffusion for high-quality large-motion vidEo frame iNterpolation. Our approach first utilizes a transformer-based tokenizer to produce refined latent representations of the intermediate frames for diffusion models. We then enhance the diffusion transformer with temporal attention across the process and incorporate a start-end frame difference embedding to guide the generation of dynamic motion. Extensive experiments demonstrate that EDEN achieves state-of-the-art results across popular benchmarks, including nearly a 10% LPIPS reduction on DAVIS and SNU-FILM, and an 8% improvement on DAIN-HD.