Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wan, Zhen, Yang, Chao-Han Huck, Tian, Jinchuan, Ye, Hanrong, Pasad, Ankita, Fu, Szu-wei, Goel, Arushi, Hachiuma, Ryo, Diao, Shizhe, Dhawan, Kunal, Ghosh, Sreyan, Hirota, Yusuke, Chen, Zhehuai, Valle, Rafael, Chu, Chenhui, Watanabe, Shinji, Wang, Yu-Chiang Frank, Ginsburg, Boris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What do Speech Foundation Models Learn? Analysis and Applications
von: Pasad, Ankita
Veröffentlicht: (2025)
von: Pasad, Ankita
Veröffentlicht: (2025)
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models
von: Wan, Zhen, et al.
Veröffentlicht: (2025)
von: Wan, Zhen, et al.
Veröffentlicht: (2025)
Training and Inference Efficiency of Encoder-Decoder Speech Models
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
von: Lu, Ke-Han, et al.
Veröffentlicht: (2024)
von: Lu, Ke-Han, et al.
Veröffentlicht: (2024)
DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
von: Lu, Ke-Han, et al.
Veröffentlicht: (2024)
von: Lu, Ke-Han, et al.
Veröffentlicht: (2024)
Self-Supervised Speech Representations are More Phonetic than Semantic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
von: Dhawan, Kunal, et al.
Veröffentlicht: (2024)
von: Dhawan, Kunal, et al.
Veröffentlicht: (2024)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
von: Xu, Hainan, et al.
Veröffentlicht: (2024)
von: Xu, Hainan, et al.
Veröffentlicht: (2024)
What Do Self-Supervised Speech Models Know About Words?
von: Pasad, Ankita, et al.
Veröffentlicht: (2023)
von: Pasad, Ankita, et al.
Veröffentlicht: (2023)
Chain-of-Thought Prompting for Speech Translation
von: Hu, Ke, et al.
Veröffentlicht: (2024)
von: Hu, Ke, et al.
Veröffentlicht: (2024)
Adapting Speech Language Model to Singing Voice Synthesis
von: Zhao, Yiwen, et al.
Veröffentlicht: (2025)
von: Zhao, Yiwen, et al.
Veröffentlicht: (2025)
UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
von: Chen, William, et al.
Veröffentlicht: (2025)
von: Chen, William, et al.
Veröffentlicht: (2025)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
von: Puvvada, Krishna C., et al.
Veröffentlicht: (2024)
von: Puvvada, Krishna C., et al.
Veröffentlicht: (2024)
Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
von: Noroozi, Vahid, et al.
Veröffentlicht: (2024)
von: Noroozi, Vahid, et al.
Veröffentlicht: (2024)
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
von: Hirota, Yusuke, et al.
Veröffentlicht: (2024)
von: Hirota, Yusuke, et al.
Veröffentlicht: (2024)
Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
von: Huang, Sung-Feng, et al.
Veröffentlicht: (2025)
von: Huang, Sung-Feng, et al.
Veröffentlicht: (2025)
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
von: Ye, Hanrong, et al.
Veröffentlicht: (2025)
von: Ye, Hanrong, et al.
Veröffentlicht: (2025)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
Multi-blank Transducers for Speech Recognition
von: Xu, Hainan, et al.
Veröffentlicht: (2022)
von: Xu, Hainan, et al.
Veröffentlicht: (2022)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
von: Fu, Szu-Wei, et al.
Veröffentlicht: (2024)
von: Fu, Szu-Wei, et al.
Veröffentlicht: (2024)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2024)
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2024)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
von: Park, Taejin, et al.
Veröffentlicht: (2024)
von: Park, Taejin, et al.
Veröffentlicht: (2024)
NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
von: Huang, He, et al.
Veröffentlicht: (2024)
von: Huang, He, et al.
Veröffentlicht: (2024)
META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
von: Wang, Jinhan, et al.
Veröffentlicht: (2024)
von: Wang, Jinhan, et al.
Veröffentlicht: (2024)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition
von: Su, Bo-Hao, et al.
Veröffentlicht: (2025)
von: Su, Bo-Hao, et al.
Veröffentlicht: (2025)
Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs
von: Langman, Ryan, et al.
Veröffentlicht: (2024)
von: Langman, Ryan, et al.
Veröffentlicht: (2024)
OMCAT: Omni Context Aware Transformer
von: Goel, Arushi, et al.
Veröffentlicht: (2024)
von: Goel, Arushi, et al.
Veröffentlicht: (2024)
OpusLM: A Family of Open Unified Speech Language Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
von: Chen, Zhehuai, et al.
Veröffentlicht: (2024)
von: Chen, Zhehuai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
What do Speech Foundation Models Learn? Analysis and Applications
von: Pasad, Ankita
Veröffentlicht: (2025) -
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
von: Peng, Yifan, et al.
Veröffentlicht: (2024) -
SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models
von: Wan, Zhen, et al.
Veröffentlicht: (2025) -
Training and Inference Efficiency of Encoder-Decoder Speech Models
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025) -
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
von: Lu, Ke-Han, et al.
Veröffentlicht: (2024)