A Deep Learning Approach for Video Coding Based on Spatial and Frequency Domain Features
作者:Chenyang Ding, Lei Chen, Shicheng Xu, Jiayi Xu · 年份:2026 · DOI:10.1109/cnml68938.2026.11453023 · 研究领域:Video Coding and Compression Technologies、Advanced Data Compression Techniques、Image and Video Quality Assessment
The latest Versatile Video Coding (VVC/H.266) standard significantly enhances compression efficiency but at the cost of drastically increased computational complexity. To address this challenge, this paper proposes a fast coding decision method based on spatial-frequency feature synergy to reduce encoding time. The method jointly extracts spatial structural features and frequency-domain transform coefficients from coding units (CUs) and employs a lightweight network for feature fusion and prediction, thereby effectively bypassing the original time-consuming exhaustive mode search. Experimental results demonstrate that the proposed method achieves a stable encoding acceleration of 21% to 28% for both CU partitioning and intra prediction tasks, while introducing only a 1.91% to 3.53% increase in bitrate. Compared to existing solutions, our approach achieves a more competitive balance between coding efficiency and computational complexity, providing an effective technical pathway for the practical application of VVC in real-time scenarios.