Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wan, Zhen, Yang, Chao-Han Huck, Tian, Jinchuan, Ye, Hanrong, Pasad, Ankita, Fu, Szu-wei, Goel, Arushi, Hachiuma, Ryo, Diao, Shizhe, Dhawan, Kunal, Ghosh, Sreyan, Hirota, Yusuke, Chen, Zhehuai, Valle, Rafael, Chu, Chenhui, Watanabe, Shinji, Wang, Yu-Chiang Frank, Ginsburg, Boris
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918508530499584
author Wan, Zhen
Yang, Chao-Han Huck
Tian, Jinchuan
Ye, Hanrong
Pasad, Ankita
Fu, Szu-wei
Goel, Arushi
Hachiuma, Ryo
Diao, Shizhe
Dhawan, Kunal
Ghosh, Sreyan
Hirota, Yusuke
Chen, Zhehuai
Valle, Rafael
Chu, Chenhui
Watanabe, Shinji
Wang, Yu-Chiang Frank
Ginsburg, Boris
author_facet Wan, Zhen
Yang, Chao-Han Huck
Tian, Jinchuan
Ye, Hanrong
Pasad, Ankita
Fu, Szu-wei
Goel, Arushi
Hachiuma, Ryo
Diao, Shizhe
Dhawan, Kunal
Ghosh, Sreyan
Hirota, Yusuke
Chen, Zhehuai
Valle, Rafael
Chu, Chenhui
Watanabe, Shinji
Wang, Yu-Chiang Frank
Ginsburg, Boris
contents We introduce a voice-agentic framework that learns one critical omni-understanding skill: knowing when to trust itself versus when to consult external audio perception. Our work is motivated by a crucial yet counterintuitive finding: naively fine-tuning an omni-model on both speech recognition and external sound understanding tasks often degrades performance, as the model can be easily misled by noisy hypotheses. To address this, our framework, Speech-Hands, recasts the problem as an explicit self-reflection decision. This learnable reflection primitive proves effective in preventing the model from being derailed by flawed external candidates. We show that this agentic action mechanism generalizes naturally from speech recognition to complex, multiple-choice audio reasoning. Across the OpenASR leaderboard, Speech-Hands consistently outperforms strong baselines by 12.1% WER on seven benchmarks. The model also achieves 77.37% accuracy and high F1 on audio QA decisions, showing robust generalization and reliability across diverse audio question answering datasets. By unifying perception and decision-making, our work offers a practical path toward more reliable and resilient audio intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2601_09413
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
Wan, Zhen
Yang, Chao-Han Huck
Tian, Jinchuan
Ye, Hanrong
Pasad, Ankita
Fu, Szu-wei
Goel, Arushi
Hachiuma, Ryo
Diao, Shizhe
Dhawan, Kunal
Ghosh, Sreyan
Hirota, Yusuke
Chen, Zhehuai
Valle, Rafael
Chu, Chenhui
Watanabe, Shinji
Wang, Yu-Chiang Frank
Ginsburg, Boris
Sound
Artificial Intelligence
Computation and Language
Multiagent Systems
Audio and Speech Processing
We introduce a voice-agentic framework that learns one critical omni-understanding skill: knowing when to trust itself versus when to consult external audio perception. Our work is motivated by a crucial yet counterintuitive finding: naively fine-tuning an omni-model on both speech recognition and external sound understanding tasks often degrades performance, as the model can be easily misled by noisy hypotheses. To address this, our framework, Speech-Hands, recasts the problem as an explicit self-reflection decision. This learnable reflection primitive proves effective in preventing the model from being derailed by flawed external candidates. We show that this agentic action mechanism generalizes naturally from speech recognition to complex, multiple-choice audio reasoning. Across the OpenASR leaderboard, Speech-Hands consistently outperforms strong baselines by 12.1% WER on seven benchmarks. The model also achieves 77.37% accuracy and high F1 on audio QA decisions, showing robust generalization and reliability across diverse audio question answering datasets. By unifying perception and decision-making, our work offers a practical path toward more reliable and resilient audio intelligence.
title Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
topic Sound
Artificial Intelligence
Computation and Language
Multiagent Systems
Audio and Speech Processing
url https://arxiv.org/abs/2601.09413