FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tu, Shuyuan, Pan, Yueming, Huang, Yinming, Han, Xintong, Xing, Zhen, Dai, Qi, Qiu, Kai, Luo, Chong, Wu, Zuxuan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918255452487680
author Tu, Shuyuan
Pan, Yueming
Huang, Yinming
Han, Xintong
Xing, Zhen
Dai, Qi
Qiu, Kai
Luo, Chong
Wu, Zuxuan
author_facet Tu, Shuyuan
Pan, Yueming
Huang, Yinming
Han, Xintong
Xing, Zhen
Dai, Qi
Qiu, Kai
Luo, Chong
Wu, Zuxuan
contents Current diffusion-based acceleration methods for long-portrait animation struggle to ensure identity (ID) consistency. This paper presents FlashPortrait, an end-to-end video diffusion transformer capable of synthesizing ID-preserving, infinite-length videos while achieving up to 6x acceleration in inference speed. In particular, FlashPortrait begins by computing the identity-agnostic facial expression features with an off-the-shelf extractor. It then introduces a Normalized Facial Expression Block to align facial features with diffusion latents by normalizing them with their respective means and variances, thereby improving identity stability in facial modeling. During inference, FlashPortrait adopts a dynamic sliding-window scheme with weighted blending in overlapping areas, ensuring smooth transitions and ID consistency in long animations. In each context window, based on the latent variation rate at particular timesteps and the derivative magnitude ratio among diffusion layers, FlashPortrait utilizes higher-order latent derivatives at the current timestep to directly predict latents at future timesteps, thereby skipping several denoising steps and achieving 6x speed acceleration. Experiments on benchmarks show the effectiveness of FlashPortrait both qualitatively and quantitatively.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16900
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction
Tu, Shuyuan
Pan, Yueming
Huang, Yinming
Han, Xintong
Xing, Zhen
Dai, Qi
Qiu, Kai
Luo, Chong
Wu, Zuxuan
Computer Vision and Pattern Recognition
Current diffusion-based acceleration methods for long-portrait animation struggle to ensure identity (ID) consistency. This paper presents FlashPortrait, an end-to-end video diffusion transformer capable of synthesizing ID-preserving, infinite-length videos while achieving up to 6x acceleration in inference speed. In particular, FlashPortrait begins by computing the identity-agnostic facial expression features with an off-the-shelf extractor. It then introduces a Normalized Facial Expression Block to align facial features with diffusion latents by normalizing them with their respective means and variances, thereby improving identity stability in facial modeling. During inference, FlashPortrait adopts a dynamic sliding-window scheme with weighted blending in overlapping areas, ensuring smooth transitions and ID consistency in long animations. In each context window, based on the latent variation rate at particular timesteps and the derivative magnitude ratio among diffusion layers, FlashPortrait utilizes higher-order latent derivatives at the current timestep to directly predict latents at future timesteps, thereby skipping several denoising steps and achieving 6x speed acceleration. Experiments on benchmarks show the effectiveness of FlashPortrait both qualitatively and quantitatively.
title FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.16900