Guardado en:
| Autores principales: | Yoo, Jaekwon, Chandiramani, Kunal, Tadimeti, Divya, Girma, Abenezer, Dhir, Chandra |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2509.04473 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Streaming Speech-to-Text Translation with a SpeechLLM
por: Parcollet, Titouan, et al.
Publicado: (2026)
por: Parcollet, Titouan, et al.
Publicado: (2026)
Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
por: Teleki, Maria, et al.
Publicado: (2025)
por: Teleki, Maria, et al.
Publicado: (2025)
Measuring the Redundancy of Decoder Layers in SpeechLLMs
por: Moumen, Adel, et al.
Publicado: (2026)
por: Moumen, Adel, et al.
Publicado: (2026)
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
por: Prakash, Jeena, et al.
Publicado: (2025)
por: Prakash, Jeena, et al.
Publicado: (2025)
Task-Lens: Cross-Task Utility Based Speech Dataset Profiling for Low-Resource Indian Languages
por: Sharma, Swati, et al.
Publicado: (2026)
por: Sharma, Swati, et al.
Publicado: (2026)
Contrastive Learning for Task-Independent SpeechLLM-Pretraining
por: Züfle, Maike, et al.
Publicado: (2024)
por: Züfle, Maike, et al.
Publicado: (2024)
Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
por: Parikh, Aditya Kamlesh, et al.
Publicado: (2026)
por: Parikh, Aditya Kamlesh, et al.
Publicado: (2026)
Slot Filling as a Reasoning Task for SpeechLLMs
por: Hacioglu, Kadri, et al.
Publicado: (2025)
por: Hacioglu, Kadri, et al.
Publicado: (2025)
Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps
por: Waldendorf, Jonas, et al.
Publicado: (2026)
por: Waldendorf, Jonas, et al.
Publicado: (2026)
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
por: Wang, Dingdong, et al.
Publicado: (2025)
por: Wang, Dingdong, et al.
Publicado: (2025)
Short-form Text Rewriting with Phi Silica
por: Tadimeti, Divya, et al.
Publicado: (2026)
por: Tadimeti, Divya, et al.
Publicado: (2026)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
por: Zhang, Linhao, et al.
Publicado: (2025)
por: Zhang, Linhao, et al.
Publicado: (2025)
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
por: Papi, Sara, et al.
Publicado: (2026)
por: Papi, Sara, et al.
Publicado: (2026)
LoASR-Bench: Evaluating Large Speech Language Models on Low-Resource Automatic Speech Recognition Across Language Families
por: Chen, Jianan, et al.
Publicado: (2026)
por: Chen, Jianan, et al.
Publicado: (2026)
SpeechPrompt: Prompting Speech Language Models for Speech Processing Tasks
por: Chang, Kai-Wei, et al.
Publicado: (2024)
por: Chang, Kai-Wei, et al.
Publicado: (2024)
SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition
por: Wu, Yihan, et al.
Publicado: (2024)
por: Wu, Yihan, et al.
Publicado: (2024)
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
por: Zhao, Mengjie, et al.
Publicado: (2026)
por: Zhao, Mengjie, et al.
Publicado: (2026)
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
por: Satish, Shree Harsha Bokkahalli, et al.
Publicado: (2025)
por: Satish, Shree Harsha Bokkahalli, et al.
Publicado: (2025)
Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages
por: Prome, Ruhina Tabasshum, et al.
Publicado: (2025)
por: Prome, Ruhina Tabasshum, et al.
Publicado: (2025)
Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks
por: Hsu, Ming-Hao, et al.
Publicado: (2023)
por: Hsu, Ming-Hao, et al.
Publicado: (2023)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
por: Song, Yuhan, et al.
Publicado: (2025)
por: Song, Yuhan, et al.
Publicado: (2025)
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
por: Elmakies, Avishai, et al.
Publicado: (2025)
por: Elmakies, Avishai, et al.
Publicado: (2025)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
por: Saengthong, Phurich, et al.
Publicado: (2025)
por: Saengthong, Phurich, et al.
Publicado: (2025)
Aligning Paralinguistic Understanding and Generation in Speech LLMs via Multi-Task Reinforcement Learning
por: Chen, Jingxiang, et al.
Publicado: (2026)
por: Chen, Jingxiang, et al.
Publicado: (2026)
Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization
por: Wadhwa, Sahil, et al.
Publicado: (2024)
por: Wadhwa, Sahil, et al.
Publicado: (2024)
Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language Models
por: Xiang, Bajian, et al.
Publicado: (2025)
por: Xiang, Bajian, et al.
Publicado: (2025)
Task Arithmetic for Language Expansion in Speech Translation
por: Cheng, Yao-Fei, et al.
Publicado: (2024)
por: Cheng, Yao-Fei, et al.
Publicado: (2024)
Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
por: Kubis, Marek, et al.
Publicado: (2025)
por: Kubis, Marek, et al.
Publicado: (2025)
Enhancing Speech Instruction Understanding and Disambiguation in Robotics via Speech Prosody
por: Sasu, David, et al.
Publicado: (2025)
por: Sasu, David, et al.
Publicado: (2025)
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
por: Yan, Canxiang, et al.
Publicado: (2025)
por: Yan, Canxiang, et al.
Publicado: (2025)
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
por: Satish, Shree Harsha Bokkahalli, et al.
Publicado: (2026)
por: Satish, Shree Harsha Bokkahalli, et al.
Publicado: (2026)
SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia
por: Ng, Ri Chi, et al.
Publicado: (2026)
por: Ng, Ri Chi, et al.
Publicado: (2026)
Can Linguistically Related Languages Guide LLM Translation in Low-Resource Settings?
por: Ramasethu, Aishwarya, et al.
Publicado: (2026)
por: Ramasethu, Aishwarya, et al.
Publicado: (2026)
Improving Semantic Understanding in Speech Language Models via Brain-tuning
por: Moussa, Omer, et al.
Publicado: (2024)
por: Moussa, Omer, et al.
Publicado: (2024)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
por: Toyin, Hawau Olamide, et al.
Publicado: (2024)
por: Toyin, Hawau Olamide, et al.
Publicado: (2024)
GHaLIB: A Multilingual Framework for Hope Speech Detection in Low-Resource Languages
por: Abdullah, Ahmed, et al.
Publicado: (2025)
por: Abdullah, Ahmed, et al.
Publicado: (2025)
SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
por: Hacioglu, Kadri, et al.
Publicado: (2025)
por: Hacioglu, Kadri, et al.
Publicado: (2025)
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
por: Fong, Seraphina, et al.
Publicado: (2025)
por: Fong, Seraphina, et al.
Publicado: (2025)
Locate-and-Focus: Enhancing Terminology Translation in Speech Language Models
por: Wu, Suhang, et al.
Publicado: (2025)
por: Wu, Suhang, et al.
Publicado: (2025)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
por: Nguyen, Tuan, et al.
Publicado: (2025)
por: Nguyen, Tuan, et al.
Publicado: (2025)
Ejemplares similares
-
Streaming Speech-to-Text Translation with a SpeechLLM
por: Parcollet, Titouan, et al.
Publicado: (2026) -
Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
por: Teleki, Maria, et al.
Publicado: (2025) -
Measuring the Redundancy of Decoder Layers in SpeechLLMs
por: Moumen, Adel, et al.
Publicado: (2026) -
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
por: Prakash, Jeena, et al.
Publicado: (2025) -
Task-Lens: Cross-Task Utility Based Speech Dataset Profiling for Low-Resource Indian Languages
por: Sharma, Swati, et al.
Publicado: (2026)