Step-Audio-R1.5 Technical Report

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Yuxin, Zhang, Xiangyu Tony, Liu, Daijiao, Tian, Fei, Deng, Yayue, Chen, Jun, Lin, Qingjian, Zhang, Haoyang, Li, Yuxin, Gong, Jinglan, Huang, Yechang, Zhao, Liang, Yao, Chengyuan, Liu, Hexin, Chng, Eng Siong, Yang, Xuerui, Yu, Gang, Zhang, Xiangyu, Jiang, Daxin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911737932939264
author Zhang, Yuxin
Zhang, Xiangyu Tony
Liu, Daijiao
Tian, Fei
Deng, Yayue
Chen, Jun
Lin, Qingjian
Zhang, Haoyang
Li, Yuxin
Gong, Jinglan
Huang, Yechang
Zhao, Liang
Yao, Chengyuan
Liu, Hexin
Chng, Eng Siong
Yang, Xuerui
Yu, Gang
Zhang, Xiangyu
Jiang, Daxin
author_facet Zhang, Yuxin
Zhang, Xiangyu Tony
Liu, Daijiao
Tian, Fei
Deng, Yayue
Chen, Jun
Lin, Qingjian
Zhang, Haoyang
Li, Yuxin
Gong, Jinglan
Huang, Yechang
Zhao, Liang
Yao, Chengyuan
Liu, Hexin
Chng, Eng Siong
Yang, Xuerui
Yu, Gang
Zhang, Xiangyu
Jiang, Daxin
contents Recent advancements in large audio language models have extended Chain-of-Thought (CoT) reasoning into the auditory domain, enabling models to tackle increasingly complex acoustic and spoken tasks. To elicit and sustain these extended reasoning chains, the prevailing paradigm -- driven by the success of text-based reasoning models -- overwhelmingly relies on Reinforcement Learning with Verified Rewards (RLVR). However, as models are strictly optimized to distill rich, continuous auditory contexts into isolated, verifiable text labels, a fundamental question arises: are we fostering true audio intelligence, or merely reducing a continuous sensory medium into a discrete puzzle? We identify this as the "verifiable reward trap." While RLVR yields remarkable scores on standardized objective benchmarks, it systematically degrades the real-world conversational feel of audio models. By prioritizing isolated correctness over acoustic nuance, RLVR reduces dynamic interactions to mechanical "answering machines," severely compromising prosodic naturalness, emotional continuity, and user immersion, particularly in long-turn dialogues. To bridge the gap between mechanical objective verification and genuine sensory empathy, we introduce Step-Audio-R1.5, marking a paradigm shift toward Reinforcement Learning from Human Feedback (RLHF) in audio reasoning. Comprehensive evaluations demonstrate that Step-Audio-R1.5 not only maintains robust analytical reasoning but profoundly transforms the interactive experience, redefining the boundaries of deeply immersive long-turn spoken dialogue.
format Preprint
id arxiv_https___arxiv_org_abs_2604_25719
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Step-Audio-R1.5 Technical Report
Zhang, Yuxin
Zhang, Xiangyu Tony
Liu, Daijiao
Tian, Fei
Deng, Yayue
Chen, Jun
Lin, Qingjian
Zhang, Haoyang
Li, Yuxin
Gong, Jinglan
Huang, Yechang
Zhao, Liang
Yao, Chengyuan
Liu, Hexin
Chng, Eng Siong
Yang, Xuerui
Yu, Gang
Zhang, Xiangyu
Jiang, Daxin
Audio and Speech Processing
Recent advancements in large audio language models have extended Chain-of-Thought (CoT) reasoning into the auditory domain, enabling models to tackle increasingly complex acoustic and spoken tasks. To elicit and sustain these extended reasoning chains, the prevailing paradigm -- driven by the success of text-based reasoning models -- overwhelmingly relies on Reinforcement Learning with Verified Rewards (RLVR). However, as models are strictly optimized to distill rich, continuous auditory contexts into isolated, verifiable text labels, a fundamental question arises: are we fostering true audio intelligence, or merely reducing a continuous sensory medium into a discrete puzzle? We identify this as the "verifiable reward trap." While RLVR yields remarkable scores on standardized objective benchmarks, it systematically degrades the real-world conversational feel of audio models. By prioritizing isolated correctness over acoustic nuance, RLVR reduces dynamic interactions to mechanical "answering machines," severely compromising prosodic naturalness, emotional continuity, and user immersion, particularly in long-turn dialogues. To bridge the gap between mechanical objective verification and genuine sensory empathy, we introduce Step-Audio-R1.5, marking a paradigm shift toward Reinforcement Learning from Human Feedback (RLHF) in audio reasoning. Comprehensive evaluations demonstrate that Step-Audio-R1.5 not only maintains robust analytical reasoning but profoundly transforms the interactive experience, redefining the boundaries of deeply immersive long-turn spoken dialogue.
title Step-Audio-R1.5 Technical Report
topic Audio and Speech Processing
url https://arxiv.org/abs/2604.25719