Audio-Aware Large Language Models as Judges for Speaking Styles

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chiang, Cheng-Han, Wang, Xiaofei, Lin, Chung-Ching, Lin, Kevin, Li, Linjie, Kopetz, Radu, Qian, Yao, Wang, Zhendong, Yang, Zhengyuan, Lee, Hung-yi, Wang, Lijuan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908396302630912
author Chiang, Cheng-Han
Wang, Xiaofei
Lin, Chung-Ching
Lin, Kevin
Li, Linjie
Kopetz, Radu
Qian, Yao
Wang, Zhendong
Yang, Zhengyuan
Lee, Hung-yi
Wang, Lijuan
author_facet Chiang, Cheng-Han
Wang, Xiaofei
Lin, Chung-Ching
Lin, Kevin
Li, Linjie
Kopetz, Radu
Qian, Yao
Wang, Zhendong
Yang, Zhengyuan
Lee, Hung-yi
Wang, Lijuan
contents Audio-aware large language models (ALLMs) can understand the textual and non-textual information in the audio input. In this paper, we explore using ALLMs as an automatic judge to assess the speaking styles of speeches. We use ALLM judges to evaluate the speeches generated by SLMs on two tasks: voice style instruction following and role-playing. The speaking style we consider includes emotion, volume, speaking pace, word emphasis, pitch control, and non-verbal elements. We use four spoken language models (SLMs) to complete the two tasks and use humans and ALLMs to judge the SLMs' responses. We compare two ALLM judges, GPT-4o-audio and Gemini-2.5-pro, with human evaluation results and show that the agreement between Gemini and human judges is comparable to the agreement between human evaluators. These promising results show that ALLMs can be used as a judge to evaluate SLMs. Our results also reveal that current SLMs, even GPT-4o-audio, still have room for improvement in controlling the speaking style and generating natural dialogues.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05984
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Audio-Aware Large Language Models as Judges for Speaking Styles
Chiang, Cheng-Han
Wang, Xiaofei
Lin, Chung-Ching
Lin, Kevin
Li, Linjie
Kopetz, Radu
Qian, Yao
Wang, Zhendong
Yang, Zhengyuan
Lee, Hung-yi
Wang, Lijuan
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Audio-aware large language models (ALLMs) can understand the textual and non-textual information in the audio input. In this paper, we explore using ALLMs as an automatic judge to assess the speaking styles of speeches. We use ALLM judges to evaluate the speeches generated by SLMs on two tasks: voice style instruction following and role-playing. The speaking style we consider includes emotion, volume, speaking pace, word emphasis, pitch control, and non-verbal elements. We use four spoken language models (SLMs) to complete the two tasks and use humans and ALLMs to judge the SLMs' responses. We compare two ALLM judges, GPT-4o-audio and Gemini-2.5-pro, with human evaluation results and show that the agreement between Gemini and human judges is comparable to the agreement between human evaluators. These promising results show that ALLMs can be used as a judge to evaluate SLMs. Our results also reveal that current SLMs, even GPT-4o-audio, still have room for improvement in controlling the speaking style and generating natural dialogues.
title Audio-Aware Large Language Models as Judges for Speaking Styles
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.05984