Optimizing Speech Language Models for Acoustic Consistency
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rohanian, Morteza, Krauthammer, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Interleaved Speech Modeling through Knowledge Distillation
von: Nouriborji, Mohammadmahdi, et al.
Veröffentlicht: (2025)
von: Nouriborji, Mohammadmahdi, et al.
Veröffentlicht: (2025)
Towards Scalable and Cross-Lingual Specialist Language Models for Oncology
von: Rohanian, Morteza, et al.
Veröffentlicht: (2025)
von: Rohanian, Morteza, et al.
Veröffentlicht: (2025)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
von: Zhao, Mengjie, et al.
Veröffentlicht: (2026)
von: Zhao, Mengjie, et al.
Veröffentlicht: (2026)
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Leveraging Zipformer Model for Effective Language Identification in Code-Switched Child-Directed Speech
von: Shankar, Lavanya, et al.
Veröffentlicht: (2025)
von: Shankar, Lavanya, et al.
Veröffentlicht: (2025)
End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
Scaling Analysis of Interleaved Speech-Text Language Models
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
Salmon: A Suite for Acoustic Language Model Evaluation
von: Maimon, Gallil, et al.
Veröffentlicht: (2024)
von: Maimon, Gallil, et al.
Veröffentlicht: (2024)
Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context
von: Ao, Junyi, et al.
Veröffentlicht: (2025)
von: Ao, Junyi, et al.
Veröffentlicht: (2025)
Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages
von: Bhogale, Kaushal Santosh, et al.
Veröffentlicht: (2026)
von: Bhogale, Kaushal Santosh, et al.
Veröffentlicht: (2026)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs
von: Quang, Trung Nguyen, et al.
Veröffentlicht: (2026)
von: Quang, Trung Nguyen, et al.
Veröffentlicht: (2026)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
Acoustic Model Optimization over Multiple Data Sources: Merging and Valuation
von: Wei, Victor Junqiu, et al.
Veröffentlicht: (2024)
von: Wei, Victor Junqiu, et al.
Veröffentlicht: (2024)
Efficient Training for Cross-lingual Speech Language Models
von: Zhou, Yan, et al.
Veröffentlicht: (2026)
von: Zhou, Yan, et al.
Veröffentlicht: (2026)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
PRiSM: Benchmarking Phone Realization in Speech Models
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
SpeechT: Findings of the First Mentorship in Speech Translation
von: Moslem, Yasmin, et al.
Veröffentlicht: (2025)
von: Moslem, Yasmin, et al.
Veröffentlicht: (2025)
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025)
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
Acoustic to Articulatory Inversion of Speech; Data Driven Approaches, Challenges, Applications, and Future Scope
von: Pillai, Leena G, et al.
Veröffentlicht: (2025)
von: Pillai, Leena G, et al.
Veröffentlicht: (2025)
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
von: Shu, Yuchun, et al.
Veröffentlicht: (2024)
von: Shu, Yuchun, et al.
Veröffentlicht: (2024)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
LAST: Language Model Aware Speech Tokenization
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
von: Eren, Eray, et al.
Veröffentlicht: (2025)
von: Eren, Eray, et al.
Veröffentlicht: (2025)
Uncertainty Modeling in Multimodal Speech Analysis Across the Psychosis Spectrum
von: Rohanian, Morteza, et al.
Veröffentlicht: (2025)
von: Rohanian, Morteza, et al.
Veröffentlicht: (2025)
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
von: Li, Xuanchen, et al.
Veröffentlicht: (2025)
von: Li, Xuanchen, et al.
Veröffentlicht: (2025)
Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation
von: He, Jiaxu, et al.
Veröffentlicht: (2026)
von: He, Jiaxu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Efficient Interleaved Speech Modeling through Knowledge Distillation
von: Nouriborji, Mohammadmahdi, et al.
Veröffentlicht: (2025) -
Towards Scalable and Cross-Lingual Specialist Language Models for Oncology
von: Rohanian, Morteza, et al.
Veröffentlicht: (2025) -
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
von: Liu, Rui, et al.
Veröffentlicht: (2024) -
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024) -
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
von: Zhao, Mengjie, et al.
Veröffentlicht: (2026)