Vision-Integrated High-Quality Neural Speech Coding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Guo, Yao, Ai, Yang, Zheng, Rui-Chen, Du, Hui-Peng, Jiang, Xiao-Hang, Ling, Zhen-Hua
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915312531668992
author Guo, Yao
Ai, Yang
Zheng, Rui-Chen
Du, Hui-Peng
Jiang, Xiao-Hang
Ling, Zhen-Hua
author_facet Guo, Yao
Ai, Yang
Zheng, Rui-Chen
Du, Hui-Peng
Jiang, Xiao-Hang
Ling, Zhen-Hua
contents This paper proposes a novel vision-integrated neural speech codec (VNSC), which aims to enhance speech coding quality by leveraging visual modality information. In VNSC, the image analysis-synthesis module extracts visual features from lip images, while the feature fusion module facilitates interaction between the image analysis-synthesis module and the speech coding module, transmitting visual information to assist the speech coding process. Depending on whether visual information is available during the inference stage, the feature fusion module integrates visual features into the speech coding module using either explicit integration or implicit distillation strategies. Experimental results confirm that integrating visual information effectively improves the quality of the decoded speech and enhances the noise robustness of the neural speech codec, without increasing the bitrate.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23379
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision-Integrated High-Quality Neural Speech Coding
Guo, Yao
Ai, Yang
Zheng, Rui-Chen
Du, Hui-Peng
Jiang, Xiao-Hang
Ling, Zhen-Hua
Audio and Speech Processing
Sound
This paper proposes a novel vision-integrated neural speech codec (VNSC), which aims to enhance speech coding quality by leveraging visual modality information. In VNSC, the image analysis-synthesis module extracts visual features from lip images, while the feature fusion module facilitates interaction between the image analysis-synthesis module and the speech coding module, transmitting visual information to assist the speech coding process. Depending on whether visual information is available during the inference stage, the feature fusion module integrates visual features into the speech coding module using either explicit integration or implicit distillation strategies. Experimental results confirm that integrating visual information effectively improves the quality of the decoded speech and enhances the noise robustness of the neural speech codec, without increasing the bitrate.
title Vision-Integrated High-Quality Neural Speech Coding
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2505.23379