Scholay

学术搜索 · AI 审稿 · LaTeX 协作

High Fidelity Video Prediction with Large Stochastic Recurrent Neural\n Networks

作者:Ruben Villegas, Arkanath Pathak, Harini Kannan, Dumitru Erhan, Quoc V. Le, Honglak Lee · 发表于:arXiv (Cornell University) · 年份:2019 · DOI:10.48550/arxiv.1911.01655 · 被引用次数:73 · 研究领域:Image and Video Quality Assessment、Image and Signal Denoising Methods、Advanced Image Processing Techniques

Predicting future video frames is extremely challenging, as there are many\nfactors of variation that make up the dynamics of how frames change through\ntime. Previously proposed solutions require complex inductive biases inside\nnetwork architectures with highly specialized computation, including\nsegmentation masks, optical flow, and foreground and background separation. In\nthis work, we question if such handcrafted architectures are necessary and\ninstead propose a different approach: finding minimal inductive bias for video\nprediction while maximizing network capacity. We investigate this question by\nperforming the first large-scale empirical study and demonstrate\nstate-of-the-art performance by learning large models on three different\ndatasets: one for modeling object interactions, one for modeling human motion,\nand one for modeling car driving.\n