A Novel Semantic Compression Approach for Ultra-low Bandwidth Voice Communication

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Collette, Ryan, Greenwood, Ross, Nicoll, Serena
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915502587117568
author Collette, Ryan
Greenwood, Ross
Nicoll, Serena
author_facet Collette, Ryan
Greenwood, Ross
Nicoll, Serena
contents While existing speech audio codecs designed for compression exploit limited forms of temporal redundancy and allow for multi-scale representations, they tend to represent all features of audio in the same way. In contrast, generative voice models designed for text-to-speech and voice transfer tasks have recently proved effective at factorizing audio signals into high-level semantic representations of fundamentally distinct features. In this paper, we leverage such representations in a novel semantic communications approach to achieve lower bitrates without sacrificing perceptual quality or suitability for specific downstream tasks. Our technique matches or outperforms existing audio codecs on transcription, sentiment analysis, and speaker verification when encoding at 2-4x lower bitrate -- notably surpassing Encodec in perceptual quality and speaker verification while using up to 4x less bitrate.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15462
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Novel Semantic Compression Approach for Ultra-low Bandwidth Voice Communication
Collette, Ryan
Greenwood, Ross
Nicoll, Serena
Sound
Audio and Speech Processing
While existing speech audio codecs designed for compression exploit limited forms of temporal redundancy and allow for multi-scale representations, they tend to represent all features of audio in the same way. In contrast, generative voice models designed for text-to-speech and voice transfer tasks have recently proved effective at factorizing audio signals into high-level semantic representations of fundamentally distinct features. In this paper, we leverage such representations in a novel semantic communications approach to achieve lower bitrates without sacrificing perceptual quality or suitability for specific downstream tasks. Our technique matches or outperforms existing audio codecs on transcription, sentiment analysis, and speaker verification when encoding at 2-4x lower bitrate -- notably surpassing Encodec in perceptual quality and speaker verification while using up to 4x less bitrate.
title A Novel Semantic Compression Approach for Ultra-low Bandwidth Voice Communication
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.15462