Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Jiamin, Lin, Ju, Huang, Yiteng, Vuong, Tyler, Lin, Zhaojiang, Yang, Zhaojun, Su, Peng, Rawat, Prashant, Srivastava, Sangeeta, Sun, Ming, Metze, Florian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
von: Lin, Ju, et al.
Veröffentlicht: (2024)
von: Lin, Ju, et al.
Veröffentlicht: (2024)
Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Effective Integration of KAN for Keyword Spotting
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
UniTalker: Conversational Speech-Visual Synthesis
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
MASV: Speaker Verification with Global and Local Context Mamba
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
Speech Recognition for Analysis of Police Radio Communication
von: Srivastava, Tejes, et al.
Veröffentlicht: (2024)
von: Srivastava, Tejes, et al.
Veröffentlicht: (2024)
MixRep: Hidden Representation Mixup for Low-Resource Speech Recognition
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
von: Yang, Mu, et al.
Veröffentlicht: (2025)
von: Yang, Mu, et al.
Veröffentlicht: (2025)
Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
von: Alavilli, Sagarika, et al.
Veröffentlicht: (2024)
von: Alavilli, Sagarika, et al.
Veröffentlicht: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
End-to-End DOA-Guided Speech Extraction in Noisy Multi-Talker Scenarios
von: Jing, Kangqi, et al.
Veröffentlicht: (2025)
von: Jing, Kangqi, et al.
Veröffentlicht: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
DEFORMER: Coupling Deformed Localized Patterns with Global Context for Robust End-to-end Speech Recognition
von: Xie, Jiamin, et al.
Veröffentlicht: (2022)
von: Xie, Jiamin, et al.
Veröffentlicht: (2022)
Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
Direct Preference Optimization for Speech Autoregressive Diffusion Models
von: Liu, Zhijun, et al.
Veröffentlicht: (2025)
von: Liu, Zhijun, et al.
Veröffentlicht: (2025)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
von: Huang, Chien-yu, et al.
Veröffentlicht: (2024)
von: Huang, Chien-yu, et al.
Veröffentlicht: (2024)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
von: Lin, Zhaofeng, et al.
Veröffentlicht: (2024)
von: Lin, Zhaofeng, et al.
Veröffentlicht: (2024)
Direct Speech to Speech Translation: A Review
von: Sarim, Mohammad, et al.
Veröffentlicht: (2025)
von: Sarim, Mohammad, et al.
Veröffentlicht: (2025)
FastTalker: Jointly Generating Speech and Conversational Gestures from Text
von: Guo, Zixin, et al.
Veröffentlicht: (2024)
von: Guo, Zixin, et al.
Veröffentlicht: (2024)
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition
von: Su, Bo-Hao, et al.
Veröffentlicht: (2025)
von: Su, Bo-Hao, et al.
Veröffentlicht: (2025)
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
von: de Groot, Dimme, et al.
Veröffentlicht: (2026)
von: de Groot, Dimme, et al.
Veröffentlicht: (2026)
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
von: Cui, Mingyu, et al.
Veröffentlicht: (2025)
von: Cui, Mingyu, et al.
Veröffentlicht: (2025)
SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition Via CNN-Transformer and Multidimensional Attention Mechanism
von: Tang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Tang, Xiaoyu, et al.
Veröffentlicht: (2024)
All Neural Low-latency Directional Speech Extraction
von: Pandey, Ashutosh, et al.
Veröffentlicht: (2024)
von: Pandey, Ashutosh, et al.
Veröffentlicht: (2024)
Fairness of Automatic Speech Recognition in Cleft Lip and Palate Speech
von: Bhattacharjee, Susmita, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Susmita, et al.
Veröffentlicht: (2025)
Aligning Speech to Languages to Enhance Code-switching Speech Recognition
von: Liu, Hexin, et al.
Veröffentlicht: (2024)
von: Liu, Hexin, et al.
Veröffentlicht: (2024)
Adaptive Knowledge Distillation for Device-Directed Speech Detection
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
von: Lin, Ju, et al.
Veröffentlicht: (2024) -
Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
von: Yang, Yufeng, et al.
Veröffentlicht: (2025) -
Directional Source Separation for Robust Speech Recognition on Smart Glasses
von: Feng, Tiantian, et al.
Veröffentlicht: (2023) -
Effective Integration of KAN for Keyword Spotting
von: Xu, Anfeng, et al.
Veröffentlicht: (2024) -
UniTalker: Conversational Speech-Visual Synthesis
von: Hu, Yifan, et al.
Veröffentlicht: (2025)