CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Kangyi, Li, Pengna, Fu, Jingwen, Wu, Yang, Liu, Yuhan, Zhou, Sanping, Wang, Jinjun
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916904615018496
author Wu, Kangyi
Li, Pengna
Fu, Jingwen
Wu, Yang
Liu, Yuhan
Zhou, Sanping
Wang, Jinjun
author_facet Wu, Kangyi
Li, Pengna
Fu, Jingwen
Wu, Yang
Liu, Yuhan
Zhou, Sanping
Wang, Jinjun
contents Emotional talking face generation aims to animate a human face in given reference images and generate a talking video that matches the content and emotion of driving audio. However, existing methods neglect that reference images may have a strong emotion that conflicts with the audio emotion, leading to severe emotion inaccuracy and distorted generated results. To tackle the issue, we introduce a cross-emotion memory network(CEM-Net), designed to generate emotional talking faces aligned with the driving audio when reference images exhibit strong emotion. Specifically, an Audio Emotion Enhancement module(AEE) is first devised with the cross-reconstruction training strategy to enhance audio emotion, overcoming the disruption from reference image emotion. Secondly, since reference images cannot provide sufficient facial motion information of the speaker under audio emotion, an Emotion Bridging Memory module(EBM) is utilized to compensate for the lacked information. It brings in expression displacement from the reference image emotion to the audio emotion and stores it in the memory.Given a cross-emotion feature as a query, the matching displacement can be retrieved at inference time. Extensive experiments have demonstrated that our CEM-Net can synthesize expressive, natural and lip-synced talking face videos with better emotion accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12368
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
Wu, Kangyi
Li, Pengna
Fu, Jingwen
Wu, Yang
Liu, Yuhan
Zhou, Sanping
Wang, Jinjun
Multimedia
Sound
Emotional talking face generation aims to animate a human face in given reference images and generate a talking video that matches the content and emotion of driving audio. However, existing methods neglect that reference images may have a strong emotion that conflicts with the audio emotion, leading to severe emotion inaccuracy and distorted generated results. To tackle the issue, we introduce a cross-emotion memory network(CEM-Net), designed to generate emotional talking faces aligned with the driving audio when reference images exhibit strong emotion. Specifically, an Audio Emotion Enhancement module(AEE) is first devised with the cross-reconstruction training strategy to enhance audio emotion, overcoming the disruption from reference image emotion. Secondly, since reference images cannot provide sufficient facial motion information of the speaker under audio emotion, an Emotion Bridging Memory module(EBM) is utilized to compensate for the lacked information. It brings in expression displacement from the reference image emotion to the audio emotion and stores it in the memory.Given a cross-emotion feature as a query, the matching displacement can be retrieved at inference time. Extensive experiments have demonstrated that our CEM-Net can synthesize expressive, natural and lip-synced talking face videos with better emotion accuracy.
title CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
topic Multimedia
Sound
url https://arxiv.org/abs/2508.12368