A Comprehensive Taxonomy and Analysis of Talking Head Synthesis: Techniques for Portrait Generation, Driving Mechanisms, and Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meng, Ming, Zhao, Yufei, Zhang, Bo, Zhu, Yonggui, Shi, Weimin, Wen, Maxwell, Fan, Zhaoxin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910491621720064
author Meng, Ming
Zhao, Yufei
Zhang, Bo
Zhu, Yonggui
Shi, Weimin
Wen, Maxwell
Fan, Zhaoxin
author_facet Meng, Ming
Zhao, Yufei
Zhang, Bo
Zhu, Yonggui
Shi, Weimin
Wen, Maxwell
Fan, Zhaoxin
contents Talking head synthesis, an advanced method for generating portrait videos from a still image driven by specific content, has garnered widespread attention in virtual reality, augmented reality and game production. Recently, significant breakthroughs have been made with the introduction of novel models such as the transformer and the diffusion model. Current methods can not only generate new content but also edit the generated material. This survey systematically reviews the technology, categorizing it into three pivotal domains: portrait generation, driven mechanisms, and editing techniques. We summarize milestone studies and critically analyze their innovations and shortcomings within each domain. Additionally, we organize an extensive collection of datasets and provide a thorough performance analysis of current methodologies based on various evaluation metrics, aiming to furnish a clear framework and robust data support for future research. Finally, we explore application scenarios of talking head synthesis, illustrate them with specific cases, and examine potential future directions.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10553
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Comprehensive Taxonomy and Analysis of Talking Head Synthesis: Techniques for Portrait Generation, Driving Mechanisms, and Editing
Meng, Ming
Zhao, Yufei
Zhang, Bo
Zhu, Yonggui
Shi, Weimin
Wen, Maxwell
Fan, Zhaoxin
Computer Vision and Pattern Recognition
Talking head synthesis, an advanced method for generating portrait videos from a still image driven by specific content, has garnered widespread attention in virtual reality, augmented reality and game production. Recently, significant breakthroughs have been made with the introduction of novel models such as the transformer and the diffusion model. Current methods can not only generate new content but also edit the generated material. This survey systematically reviews the technology, categorizing it into three pivotal domains: portrait generation, driven mechanisms, and editing techniques. We summarize milestone studies and critically analyze their innovations and shortcomings within each domain. Additionally, we organize an extensive collection of datasets and provide a thorough performance analysis of current methodologies based on various evaluation metrics, aiming to furnish a clear framework and robust data support for future research. Finally, we explore application scenarios of talking head synthesis, illustrate them with specific cases, and examine potential future directions.
title A Comprehensive Taxonomy and Analysis of Talking Head Synthesis: Techniques for Portrait Generation, Driving Mechanisms, and Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.10553