Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yingjie, Cao, Jiezhang, Zhang, Zicheng, Wen, Farong, Jiang, Yanwei, Jia, Jun, Liu, Xiaohong, Min, Xiongkuo, Zhai, Guangtao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916872786542592
author Zhou, Yingjie
Cao, Jiezhang
Zhang, Zicheng
Wen, Farong
Jiang, Yanwei
Jia, Jun
Liu, Xiaohong
Min, Xiongkuo
Zhai, Guangtao
author_facet Zhou, Yingjie
Cao, Jiezhang
Zhang, Zicheng
Wen, Farong
Jiang, Yanwei
Jia, Jun
Liu, Xiaohong
Min, Xiongkuo
Zhai, Guangtao
contents Speech-driven methods for portraits are figuratively known as "Talkers" because of their capability to synthesize speaking mouth shapes and facial movements. Especially with the rapid development of the Text-to-Image (T2I) models, AI-Generated Talking Heads (AGTHs) have gradually become an emerging digital human media. However, challenges persist regarding the quality of these talkers and AGTHs they generate, and comprehensive studies addressing these issues remain limited. To address this gap, this paper presents the largest AGTH quality assessment dataset THQA-10K to date, which selects 12 prominent T2I models and 14 advanced talkers to generate AGTHs for 14 prompts. After excluding instances where AGTH generation is unsuccessful, the THQA-10K dataset contains 10,457 AGTHs. Then, volunteers are recruited to subjectively rate the AGTHs and give the corresponding distortion categories. In our analysis for subjective experimental results, we evaluate the performance of talkers in terms of generalizability and quality, and also expose the distortions of existing AGTHs. Finally, an objective quality assessment method based on the first frame, Y-T slice and tone-lip consistency is proposed. Experimental results show that this method can achieve state-of-the-art (SOTA) performance in AGTH quality assessment. The work is released at https://github.com/zyj-2000/Talker.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23343
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
Zhou, Yingjie
Cao, Jiezhang
Zhang, Zicheng
Wen, Farong
Jiang, Yanwei
Jia, Jun
Liu, Xiaohong
Min, Xiongkuo
Zhai, Guangtao
Computer Vision and Pattern Recognition
Image and Video Processing
Speech-driven methods for portraits are figuratively known as "Talkers" because of their capability to synthesize speaking mouth shapes and facial movements. Especially with the rapid development of the Text-to-Image (T2I) models, AI-Generated Talking Heads (AGTHs) have gradually become an emerging digital human media. However, challenges persist regarding the quality of these talkers and AGTHs they generate, and comprehensive studies addressing these issues remain limited. To address this gap, this paper presents the largest AGTH quality assessment dataset THQA-10K to date, which selects 12 prominent T2I models and 14 advanced talkers to generate AGTHs for 14 prompts. After excluding instances where AGTH generation is unsuccessful, the THQA-10K dataset contains 10,457 AGTHs. Then, volunteers are recruited to subjectively rate the AGTHs and give the corresponding distortion categories. In our analysis for subjective experimental results, we evaluate the performance of talkers in terms of generalizability and quality, and also expose the distortions of existing AGTHs. Finally, an objective quality assessment method based on the first frame, Y-T slice and tone-lip consistency is proposed. Experimental results show that this method can achieve state-of-the-art (SOTA) performance in AGTH quality assessment. The work is released at https://github.com/zyj-2000/Talker.
title Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2507.23343