Data-driven Head Motion Generation through Natural Gaze-Head Coordination

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Xiaohan, Wen, Yilin, Sugano, Yusuke
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914599445463040
author Liu, Xiaohan
Wen, Yilin
Sugano, Yusuke
author_facet Liu, Xiaohan
Wen, Yilin
Sugano, Yusuke
contents We present the first data-driven approach to model temporal gaze-head coordination from large-scale in-the-wild facial videos. To obtain training data for generalizable learning, we propose an automatic pipeline that extracts natural yet diverse gaze and head motions with off-the-shelf appearance-based gaze estimators. To capture the probabilistic correlation and temporal dynamics of gaze-head coordination, we build our model on a generative conditional Variational Autoencoder for plausible yet diverse gaze-conditioned head motion generations. We further apply our framework to gaze-controlled facial video generation, where we enable video generation with natural and realistic head motion correlated to the input gaze - an aspect that has not been emphasized before. Human evaluation and quantitative comparisons demonstrate our method's effectiveness and validate our design choices, with evaluators showing statistically significant preference for our approach over baseline methods.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25810
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Data-driven Head Motion Generation through Natural Gaze-Head Coordination
Liu, Xiaohan
Wen, Yilin
Sugano, Yusuke
Computer Vision and Pattern Recognition
We present the first data-driven approach to model temporal gaze-head coordination from large-scale in-the-wild facial videos. To obtain training data for generalizable learning, we propose an automatic pipeline that extracts natural yet diverse gaze and head motions with off-the-shelf appearance-based gaze estimators. To capture the probabilistic correlation and temporal dynamics of gaze-head coordination, we build our model on a generative conditional Variational Autoencoder for plausible yet diverse gaze-conditioned head motion generations. We further apply our framework to gaze-controlled facial video generation, where we enable video generation with natural and realistic head motion correlated to the input gaze - an aspect that has not been emphasized before. Human evaluation and quantitative comparisons demonstrate our method's effectiveness and validate our design choices, with evaluators showing statistically significant preference for our approach over baseline methods.
title Data-driven Head Motion Generation through Natural Gaze-Head Coordination
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.25810