Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhan, Jun, Han, Mingyang, Xie, Yuxuan, Wang, Chen, Zhang, Dong, Huang, Kexin, Shi, Haoxiang, Wang, DongXiao, Song, Tengtao, Cheng, Qinyuan, Li, Shimin, Song, Jun, Qiu, Xipeng, Zheng, Bo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2509.09716
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916959697764352
author Zhan, Jun
Han, Mingyang
Xie, Yuxuan
Wang, Chen
Zhang, Dong
Huang, Kexin
Shi, Haoxiang
Wang, DongXiao
Song, Tengtao
Cheng, Qinyuan
Li, Shimin
Song, Jun
Qiu, Xipeng
Zheng, Bo
author_facet Zhan, Jun
Han, Mingyang
Xie, Yuxuan
Wang, Chen
Zhang, Dong
Huang, Kexin
Shi, Haoxiang
Wang, DongXiao
Song, Tengtao
Cheng, Qinyuan
Li, Shimin
Song, Jun
Qiu, Xipeng
Zheng, Bo
contents Spoken language models (SLMs) have emerged as a unified paradigm for speech understanding and generation, enabling natural human machine interaction. However, while most progress has focused on semantic accuracy and instruction following, the ability of SLMs to adapt their speaking style based on spoken instructions has received limited attention. We introduce Voice Style Adaptation (VSA), a new task that examines whether SLMs can modify their speaking style, such as timbre, prosody, or persona following natural language spoken commands. To study this task, we present VStyle, a bilingual (Chinese & English) benchmark covering four categories of speech generation: acoustic attributes, natural language instruction, role play, and implicit empathy. We also introduce the Large Audio Language Model as a Judge (LALM as a Judge) framework, which progressively evaluates outputs along textual faithfulness, style adherence, and naturalness, ensuring reproducible and objective assessment. Experiments on commercial systems and open source SLMs demonstrate that current models face clear limitations in controllable style adaptation, highlighting both the novelty and challenge of this task. By releasing VStyle and its evaluation toolkit, we aim to provide the community with a foundation for advancing human centered spoken interaction. The dataset and code are publicly available at \href{https://junzhan2000.github.io/VStyle.github.io/}{project's homepage}.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09716
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
Zhan, Jun
Han, Mingyang
Xie, Yuxuan
Wang, Chen
Zhang, Dong
Huang, Kexin
Shi, Haoxiang
Wang, DongXiao
Song, Tengtao
Cheng, Qinyuan
Li, Shimin
Song, Jun
Qiu, Xipeng
Zheng, Bo
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
Spoken language models (SLMs) have emerged as a unified paradigm for speech understanding and generation, enabling natural human machine interaction. However, while most progress has focused on semantic accuracy and instruction following, the ability of SLMs to adapt their speaking style based on spoken instructions has received limited attention. We introduce Voice Style Adaptation (VSA), a new task that examines whether SLMs can modify their speaking style, such as timbre, prosody, or persona following natural language spoken commands. To study this task, we present VStyle, a bilingual (Chinese & English) benchmark covering four categories of speech generation: acoustic attributes, natural language instruction, role play, and implicit empathy. We also introduce the Large Audio Language Model as a Judge (LALM as a Judge) framework, which progressively evaluates outputs along textual faithfulness, style adherence, and naturalness, ensuring reproducible and objective assessment. Experiments on commercial systems and open source SLMs demonstrate that current models face clear limitations in controllable style adaptation, highlighting both the novelty and challenge of this task. By releasing VStyle and its evaluation toolkit, we aim to provide the community with a foundation for advancing human centered spoken interaction. The dataset and code are publicly available at \href{https://junzhan2000.github.io/VStyle.github.io/}{project's homepage}.
title VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2509.09716