A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head Videos
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910361857294336 |
|---|---|
| author | Zhang, Weixia Zhu, Chengguang Gao, Jingnan Yan, Yichao Zhai, Guangtao Yang, Xiaokang |
| author_facet | Zhang, Weixia Zhu, Chengguang Gao, Jingnan Yan, Yichao Zhai, Guangtao Yang, Xiaokang |
| contents | The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation research lags behind the development of talking head generation techniques. Existing literature relies on heuristic quantitative metrics without human validation, hindering accurate progress assessment. To address this gap, we collect talking head videos generated from four generative methods and conduct controlled psychophysical experiments on visual quality, lip-audio synchronization, and head movement naturalness. Our experiments validate consistency between model predictions and human annotations, identifying metrics that align better with human opinions than widely-used measures. We believe our work will facilitate performance evaluation and model development, providing insights into AIGC in a broader context. Code and data will be made available at https://github.com/zwx8981/ADTH-QA. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_06421 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head Videos Zhang, Weixia Zhu, Chengguang Gao, Jingnan Yan, Yichao Zhai, Guangtao Yang, Xiaokang Computer Vision and Pattern Recognition The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation research lags behind the development of talking head generation techniques. Existing literature relies on heuristic quantitative metrics without human validation, hindering accurate progress assessment. To address this gap, we collect talking head videos generated from four generative methods and conduct controlled psychophysical experiments on visual quality, lip-audio synchronization, and head movement naturalness. Our experiments validate consistency between model predictions and human annotations, identifying metrics that align better with human opinions than widely-used measures. We believe our work will facilitate performance evaluation and model development, providing insights into AIGC in a broader context. Code and data will be made available at https://github.com/zwx8981/ADTH-QA. |
| title | A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head Videos |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2403.06421 |