Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation

作者:Dingcheng Zhen, Shunshun Yin, Shiyang Qin, Yi Hou, Ziwei Zhang, Siyuan Liu, Qi Gan, Ming Tao · 年份:2025 · DOI:10.1109/cvpr52734.2025.01963 · 被引用次数:3 · 研究领域:Human Motion and Animation、Advanced Vision and Imaging、3D Surveying and Cultural Heritage

In this work, we introduce the first autoregressive framework for real-time, audio-driven portrait animation, a.k.a, talking head. Beyond the challenge of lengthy animation times, a critical challenge in realistic talking head generation lies in preserving the natural movement of diverse body parts. To this end, we propose Teller, the first streaming audio-driven protrait animation framework with autoregressive motion generation. Specifically, Teller first decomposes facial and body detail animation into two components: Facial Motion Latent Generation (FMLG) based on an autoregressive transfromer, and movement authenticity refinement using a Efficient Temporal Module (ETM). Concretely, FMLG employs a Residual VQ model to map the facial motion latent from the implicit keypoint-based model into discrete motion tokens, which are then temporally sliced with audio embeddings. This enables the AR tranformer to learn real-time, stream-based mappings from audio to motion. Furthermore, Teller incorporate ETM to capture finer motion details. This module ensures the physical consistency of body parts and accessories, such as neck muscles and earrings, improving the realism of these movements. Teller is designed to be efficient, surpassing the inference speed of diffusion-based models (Hallo 20.93s vs. ${\mathbf{Teller0}}{\mathbf{.92s}}$ for one second video generation), and achieves a real-time streaming performance of up to ${\mathbf{25 FPS}}$. Extensive experiments demonstrate that ou...