Audio-Visual Cross-Modal Compression for Generative Face Video Coding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Youmin, Guo, Mengxi, Zhao, Shijie, Li, Weiqi, Li, Junlin, Zhang, Li, Zhang, Jian
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911324214132736
author Xu, Youmin
Guo, Mengxi
Zhao, Shijie
Li, Weiqi
Li, Junlin
Zhang, Li
Zhang, Jian
author_facet Xu, Youmin
Guo, Mengxi
Zhao, Shijie
Li, Weiqi
Li, Junlin
Zhang, Li
Zhang, Jian
contents Generative face video coding (GFVC) is vital for modern applications like video conferencing, yet existing methods primarily focus on video motion while neglecting the significant bitrate contribution of audio. Despite the well-established correlation between audio and lip movements, this cross-modal coherence has not been systematically exploited for compression. To address this, we propose an Audio-Visual Cross-Modal Compression (AVCC) framework that jointly compresses audio and video streams. Our framework extracts motion information from video and tokenizes audio features, then aligns them through a unified audio-video diffusion process. This allows synchronized reconstruction of both modalities from a shared representation. In extremely low-rate scenarios, AVCC can even reconstruct one modality from the other. Experiments show that AVCC significantly outperforms the Versatile Video Coding (VVC) standard and state-of-the-art GFVC schemes in rate-distortion performance, paving the way for more efficient multimodal communication systems.
format Preprint
id arxiv_https___arxiv_org_abs_2512_15262
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Audio-Visual Cross-Modal Compression for Generative Face Video Coding
Xu, Youmin
Guo, Mengxi
Zhao, Shijie
Li, Weiqi
Li, Junlin
Zhang, Li
Zhang, Jian
Image and Video Processing
Multimedia
Generative face video coding (GFVC) is vital for modern applications like video conferencing, yet existing methods primarily focus on video motion while neglecting the significant bitrate contribution of audio. Despite the well-established correlation between audio and lip movements, this cross-modal coherence has not been systematically exploited for compression. To address this, we propose an Audio-Visual Cross-Modal Compression (AVCC) framework that jointly compresses audio and video streams. Our framework extracts motion information from video and tokenizes audio features, then aligns them through a unified audio-video diffusion process. This allows synchronized reconstruction of both modalities from a shared representation. In extremely low-rate scenarios, AVCC can even reconstruct one modality from the other. Experiments show that AVCC significantly outperforms the Versatile Video Coding (VVC) standard and state-of-the-art GFVC schemes in rate-distortion performance, paving the way for more efficient multimodal communication systems.
title Audio-Visual Cross-Modal Compression for Generative Face Video Coding
topic Image and Video Processing
Multimedia
url https://arxiv.org/abs/2512.15262