A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Weixia, Zhu, Chengguang, Gao, Jingnan, Yan, Yichao, Zhai, Guangtao, Yang, Xiaokang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910361857294336
author Zhang, Weixia
Zhu, Chengguang
Gao, Jingnan
Yan, Yichao
Zhai, Guangtao
Yang, Xiaokang
author_facet Zhang, Weixia
Zhu, Chengguang
Gao, Jingnan
Yan, Yichao
Zhai, Guangtao
Yang, Xiaokang
contents The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation research lags behind the development of talking head generation techniques. Existing literature relies on heuristic quantitative metrics without human validation, hindering accurate progress assessment. To address this gap, we collect talking head videos generated from four generative methods and conduct controlled psychophysical experiments on visual quality, lip-audio synchronization, and head movement naturalness. Our experiments validate consistency between model predictions and human annotations, identifying metrics that align better with human opinions than widely-used measures. We believe our work will facilitate performance evaluation and model development, providing insights into AIGC in a broader context. Code and data will be made available at https://github.com/zwx8981/ADTH-QA.
format Preprint
id arxiv_https___arxiv_org_abs_2403_06421
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head Videos
Zhang, Weixia
Zhu, Chengguang
Gao, Jingnan
Yan, Yichao
Zhai, Guangtao
Yang, Xiaokang
Computer Vision and Pattern Recognition
The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation research lags behind the development of talking head generation techniques. Existing literature relies on heuristic quantitative metrics without human validation, hindering accurate progress assessment. To address this gap, we collect talking head videos generated from four generative methods and conduct controlled psychophysical experiments on visual quality, lip-audio synchronization, and head movement naturalness. Our experiments validate consistency between model predictions and human annotations, identifying metrics that align better with human opinions than widely-used measures. We believe our work will facilitate performance evaluation and model development, providing insights into AIGC in a broader context. Code and data will be made available at https://github.com/zwx8981/ADTH-QA.
title A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.06421