Audio-Driven Talking Face Generation with Blink Embedding and Hash Grid Landmarks Encoding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Yuhui, Yu, Hui, Liang, Wei, Zhang, Sunjie
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915757366968320
author Zhang, Yuhui
Yu, Hui
Liang, Wei
Zhang, Sunjie
author_facet Zhang, Yuhui
Yu, Hui
Liang, Wei
Zhang, Sunjie
contents Dynamic Neural Radiance Fields (NeRF) have demonstrated considerable success in generating high-fidelity 3D models of talking portraits. Despite significant advancements in the rendering speed and generation quality, challenges persist in accurately and efficiently capturing mouth movements in talking portraits. To tackle this challenge, we propose an automatic method based on blink embedding and hash grid landmarks encoding in this study, which can substantially enhance the fidelity of talking faces. Specifically, we leverage facial features encoded as conditional features and integrate audio features as residual terms into our model through a Dynamic Landmark Transformer. Furthermore, we employ neural radiance fields to model the entire face, resulting in a lifelike face representation. Experimental evaluations have validated the superiority of our approach to existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18849
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Audio-Driven Talking Face Generation with Blink Embedding and Hash Grid Landmarks Encoding
Zhang, Yuhui
Yu, Hui
Liang, Wei
Zhang, Sunjie
Computer Vision and Pattern Recognition
Dynamic Neural Radiance Fields (NeRF) have demonstrated considerable success in generating high-fidelity 3D models of talking portraits. Despite significant advancements in the rendering speed and generation quality, challenges persist in accurately and efficiently capturing mouth movements in talking portraits. To tackle this challenge, we propose an automatic method based on blink embedding and hash grid landmarks encoding in this study, which can substantially enhance the fidelity of talking faces. Specifically, we leverage facial features encoded as conditional features and integrate audio features as residual terms into our model through a Dynamic Landmark Transformer. Furthermore, we employ neural radiance fields to model the entire face, resulting in a lifelike face representation. Experimental evaluations have validated the superiority of our approach to existing methods.
title Audio-Driven Talking Face Generation with Blink Embedding and Hash Grid Landmarks Encoding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.18849