Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huo, Mingyue, Tseng, Wei-Cheng, Shao, Yiwen, Zhang, Hao, Yu, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917092110893056
author Huo, Mingyue
Tseng, Wei-Cheng
Shao, Yiwen
Zhang, Hao
Yu, Dong
author_facet Huo, Mingyue
Tseng, Wei-Cheng
Shao, Yiwen
Zhang, Hao
Yu, Dong
contents Human voice encodes both identity and paralinguistic cues, yet encoders in large audio-language models (LALMs) rarely balance both aspects. In this work, we present a study toward building a general-purpose voice encoder that captures nuanced voice cues. Through a comprehensive evaluation, we find that multi-task training yields the most balanced representations, whereas contrastive language-audio pretraining (CLAP) primarily improves retrieval without enhancing paralinguistic understanding. Our final encoder, Auden-Voice, also demonstrates strong performance when integrated with LLMs. The code and training recipes will be released with the audio understanding toolkit Auden.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15145
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
Huo, Mingyue
Tseng, Wei-Cheng
Shao, Yiwen
Zhang, Hao
Yu, Dong
Audio and Speech Processing
Sound
Human voice encodes both identity and paralinguistic cues, yet encoders in large audio-language models (LALMs) rarely balance both aspects. In this work, we present a study toward building a general-purpose voice encoder that captures nuanced voice cues. Through a comprehensive evaluation, we find that multi-task training yields the most balanced representations, whereas contrastive language-audio pretraining (CLAP) primarily improves retrieval without enhancing paralinguistic understanding. Our final encoder, Auden-Voice, also demonstrates strong performance when integrated with LLMs. The code and training recipes will be released with the audio understanding toolkit Auden.
title Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2511.15145