JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Park, Sungjoon, Park, Minsik, Lee, Haneol, Yun, Jaesub, Lee, Donggeon
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913962120970240
author Park, Sungjoon
Park, Minsik
Lee, Haneol
Yun, Jaesub
Lee, Donggeon
author_facet Park, Sungjoon
Park, Minsik
Lee, Haneol
Yun, Jaesub
Lee, Donggeon
contents In this work, we revisit the effectiveness of 3DMM for talking head synthesis by jointly learning a 3D face reconstruction model and a talking head synthesis model. This enables us to obtain a FACS-based blendshape representation of facial expressions that is optimized for talking head synthesis. This contrasts with previous methods that either fit 3DMM parameters to 2D landmarks or rely on pretrained face reconstruction models. Not only does our approach increase the quality of the generated face, but it also allows us to take advantage of the blendshape representation to modify just the mouth region for the purpose of audio-based lip-sync. To this end, we propose a novel lip-sync pipeline that, unlike previous methods, decouples the original chin contour from the lip-synced chin contour, and reduces flickering near the mouth.
format Preprint
id arxiv_https___arxiv_org_abs_2507_20452
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync
Park, Sungjoon
Park, Minsik
Lee, Haneol
Yun, Jaesub
Lee, Donggeon
Computer Vision and Pattern Recognition
In this work, we revisit the effectiveness of 3DMM for talking head synthesis by jointly learning a 3D face reconstruction model and a talking head synthesis model. This enables us to obtain a FACS-based blendshape representation of facial expressions that is optimized for talking head synthesis. This contrasts with previous methods that either fit 3DMM parameters to 2D landmarks or rely on pretrained face reconstruction models. Not only does our approach increase the quality of the generated face, but it also allows us to take advantage of the blendshape representation to modify just the mouth region for the purpose of audio-based lip-sync. To this end, we propose a novel lip-sync pipeline that, unlike previous methods, decouples the original chin contour from the lip-synced chin contour, and reduces flickering near the mouth.
title JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.20452