EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wei, Jiangchuan, Yan, Shiyue, Lin, Wenfeng, Liu, Boyuan, Chen, Renjie, Guo, Mingyu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912249416777728
author Wei, Jiangchuan
Yan, Shiyue
Lin, Wenfeng
Liu, Boyuan
Chen, Renjie
Guo, Mingyu
author_facet Wei, Jiangchuan
Yan, Shiyue
Lin, Wenfeng
Liu, Boyuan
Chen, Renjie
Guo, Mingyu
contents Recent advancements in video generation have significantly impacted various downstream applications, particularly in identity-preserving video generation (IPT2V). However, existing methods struggle with "copy-paste" artifacts and low similarity issues, primarily due to their reliance on low-level facial image information. This dependence can result in rigid facial appearances and artifacts reflecting irrelevant details. To address these challenges, we propose EchoVideo, which employs two key strategies: (1) an Identity Image-Text Fusion Module (IITF) that integrates high-level semantic features from text, capturing clean facial identity representations while discarding occlusions, poses, and lighting variations to avoid the introduction of artifacts; (2) a two-stage training strategy, incorporating a stochastic method in the second phase to randomly utilize shallow facial information. The objective is to balance the enhancements in fidelity provided by shallow features while mitigating excessive reliance on them. This strategy encourages the model to utilize high-level features during training, ultimately fostering a more robust representation of facial identities. EchoVideo effectively preserves facial identities and maintains full-body integrity. Extensive experiments demonstrate that it achieves excellent results in generating high-quality, controllability and fidelity videos.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13452
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
Wei, Jiangchuan
Yan, Shiyue
Lin, Wenfeng
Liu, Boyuan
Chen, Renjie
Guo, Mingyu
Computer Vision and Pattern Recognition
Recent advancements in video generation have significantly impacted various downstream applications, particularly in identity-preserving video generation (IPT2V). However, existing methods struggle with "copy-paste" artifacts and low similarity issues, primarily due to their reliance on low-level facial image information. This dependence can result in rigid facial appearances and artifacts reflecting irrelevant details. To address these challenges, we propose EchoVideo, which employs two key strategies: (1) an Identity Image-Text Fusion Module (IITF) that integrates high-level semantic features from text, capturing clean facial identity representations while discarding occlusions, poses, and lighting variations to avoid the introduction of artifacts; (2) a two-stage training strategy, incorporating a stochastic method in the second phase to randomly utilize shallow facial information. The objective is to balance the enhancements in fidelity provided by shallow features while mitigating excessive reliance on them. This strategy encourages the model to utilize high-level features during training, ultimately fostering a more robust representation of facial identities. EchoVideo effectively preserves facial identities and maintains full-body integrity. Extensive experiments demonstrate that it achieves excellent results in generating high-quality, controllability and fidelity videos.
title EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.13452