Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Long, Jianzhi, Sun, Wenhao, Tu, Rongcheng, Tao, Dacheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917208872976384
author Long, Jianzhi
Sun, Wenhao
Tu, Rongcheng
Tao, Dacheng
author_facet Long, Jianzhi
Sun, Wenhao
Tu, Rongcheng
Tao, Dacheng
contents Diffusion-based talking head models generate high-quality, photorealistic videos but suffer from slow inference, limiting practical applications. Existing acceleration methods for general diffusion models fail to exploit the temporal and spatial redundancies unique to talking head generation. In this paper, we propose a task-specific framework addressing these inefficiencies through two key innovations. First, we introduce Lightning-fast Caching-based Parallel denoising prediction (LightningCP), caching static features to bypass most model layers in inference time. We also enable parallel prediction using cached features and estimated noisy latents as inputs, efficiently bypassing sequential sampling. Second, we propose Decoupled Foreground Attention (DFA) to further accelerate attention computations, exploiting the spatial decoupling in talking head videos to restrict attention to dynamic foreground regions. Additionally, we remove reference features in certain layers to bring extra speedup. Extensive experiments demonstrate that our framework significantly improves inference speed while preserving video quality.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00052
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation
Long, Jianzhi
Sun, Wenhao
Tu, Rongcheng
Tao, Dacheng
Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
Diffusion-based talking head models generate high-quality, photorealistic videos but suffer from slow inference, limiting practical applications. Existing acceleration methods for general diffusion models fail to exploit the temporal and spatial redundancies unique to talking head generation. In this paper, we propose a task-specific framework addressing these inefficiencies through two key innovations. First, we introduce Lightning-fast Caching-based Parallel denoising prediction (LightningCP), caching static features to bypass most model layers in inference time. We also enable parallel prediction using cached features and estimated noisy latents as inputs, efficiently bypassing sequential sampling. Second, we propose Decoupled Foreground Attention (DFA) to further accelerate attention computations, exploiting the spatial decoupling in talking head videos to restrict attention to dynamic foreground regions. Additionally, we remove reference features in certain layers to bring extra speedup. Extensive experiments demonstrate that our framework significantly improves inference speed while preserving video quality.
title Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation
topic Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.00052