Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Divide-and-Conquer Completion Network for Video Inpainting

作者:Zhiliang Wu, Changchang Sun, Hanyu Xuan, Kang Zhang, Yan Yan · 发表于:IEEE Transactions on Circuits and Systems for Video Technology · 年份:2022 · DOI:10.1109/tcsvt.2022.3225911 · 被引用次数:34 · 研究领域:Generative Adversarial Networks and Image Synthesis、Advanced Image Processing Techniques、Advanced Vision and Imaging

Video inpainting aims to utilize plausible contents to complete missing regions in the video. For different components, the reconstruction targets of missing regions are different,e.g.,smoothness preserving for flat regions, sharpening for edges and textures. Typically, existing methods treat the missing regions as a whole and holistically train the model by optimizing homogenous pixel-wise losses (e.g.,MSE). In this way, the trained models will be easily dominated and determined by flat regions, failing to infer realistic details (edges and textures) that are difficult to reconstruct but necessary for practical applications. In this paper, we propose a divide-and-conquer completion network for video inpainting. In particular, our network first uses discrete wavelet transform to decompose the deep features into low-frequency components containing structural information (flat regions) and high-frequency components involving detailed texture information. Thereafter, we feed these components into different branches and adopt the temporal attention feature aggregation module to generate missing contents, separately. It hence can realize flexible supervision utilizing the intermediate supervision learning strategy for each component, which has not been noticed and explored by current state-of-the-art video inpainting methods. Furthermore, we adopt a gradient-weighted reconstruction loss to supervise the completed frame reconstruction process, which can use the gradients in all dir...