Efficient Training for Cross-lingual Speech Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhou, Yan, Fang, Qingkai, Hong, Yun, Feng, Yang |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
par: Zhang, Shaolei, et autres
Publié: (2025)
par: Zhang, Shaolei, et autres
Publié: (2025)
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
par: Fang, Qingkai, et autres
Publié: (2025)
par: Fang, Qingkai, et autres
Publié: (2025)
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
par: Fang, Qingkai, et autres
Publié: (2024)
par: Fang, Qingkai, et autres
Publié: (2024)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
par: Fang, Qingkai, et autres
Publié: (2024)
par: Fang, Qingkai, et autres
Publié: (2024)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
par: Zhang, Shaolei, et autres
Publié: (2024)
par: Zhang, Shaolei, et autres
Publié: (2024)
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
par: Ma, Zhengrui, et autres
Publié: (2024)
par: Ma, Zhengrui, et autres
Publié: (2024)
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
par: Fang, Qingkai, et autres
Publié: (2024)
par: Fang, Qingkai, et autres
Publié: (2024)
Cross-Attention is Half Explanation in Speech-to-Text Models
par: Papi, Sara, et autres
Publié: (2025)
par: Papi, Sara, et autres
Publié: (2025)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
par: Lin, Tzu-Quan, et autres
Publié: (2025)
par: Lin, Tzu-Quan, et autres
Publié: (2025)
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
par: Zhang, Yuhao, et autres
Publié: (2025)
par: Zhang, Yuhao, et autres
Publié: (2025)
Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation
par: Feng, Bo-Han, et autres
Publié: (2026)
par: Feng, Bo-Han, et autres
Publié: (2026)
Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
par: Shi, Jiacheng, et autres
Publié: (2025)
par: Shi, Jiacheng, et autres
Publié: (2025)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
par: Kim, Ji-Hoon, et autres
Publié: (2024)
par: Kim, Ji-Hoon, et autres
Publié: (2024)
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
par: Gaido, Marco, et autres
Publié: (2024)
par: Gaido, Marco, et autres
Publié: (2024)
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
par: Ginjala, Srishti, et autres
Publié: (2026)
par: Ginjala, Srishti, et autres
Publié: (2026)
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
par: Papi, Sara, et autres
Publié: (2026)
par: Papi, Sara, et autres
Publié: (2026)
Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment
par: Gao, Yan, et autres
Publié: (2025)
par: Gao, Yan, et autres
Publié: (2025)
Raon-Speech Technical Report
par: Kim, Beomsoo, et autres
Publié: (2026)
par: Kim, Beomsoo, et autres
Publié: (2026)
ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition
par: Lee, Junseok, et autres
Publié: (2026)
par: Lee, Junseok, et autres
Publié: (2026)
Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data
par: Kumar, Gokul Karthik, et autres
Publié: (2025)
par: Kumar, Gokul Karthik, et autres
Publié: (2025)
Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
par: Riera, Pablo, et autres
Publié: (2026)
par: Riera, Pablo, et autres
Publié: (2026)
Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC)
par: Wang, Peidong
Publié: (2026)
par: Wang, Peidong
Publié: (2026)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
par: Zhang, Xueyao, et autres
Publié: (2025)
par: Zhang, Xueyao, et autres
Publié: (2025)
XCB: an effective contextual biasing approach to bias cross-lingual phrases in speech recognition
par: Wan, Xucheng, et autres
Publié: (2024)
par: Wan, Xucheng, et autres
Publié: (2024)
SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation
par: Liu, Ruohan, et autres
Publié: (2026)
par: Liu, Ruohan, et autres
Publié: (2026)
Exploring Machine Learning and Language Models for Multimodal Depression Detection
par: Hong, Javier Si Zhao, et autres
Publié: (2025)
par: Hong, Javier Si Zhao, et autres
Publié: (2025)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
par: Zhang, Yuhao, et autres
Publié: (2025)
par: Zhang, Yuhao, et autres
Publié: (2025)
Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts
par: Jin, Hojun, et autres
Publié: (2025)
par: Jin, Hojun, et autres
Publié: (2025)
Efficient Dialect-Aware Modeling and Conditioning for Low-Resource Taiwanese Hakka Speech Processing
par: Peng, An-Ci, et autres
Publié: (2026)
par: Peng, An-Ci, et autres
Publié: (2026)
Efficient Compression of Multitask Multilingual Speech Models
par: Ferraz, Thomas Palmeira
Publié: (2024)
par: Ferraz, Thomas Palmeira
Publié: (2024)
FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
par: Papi, Sara, et autres
Publié: (2025)
par: Papi, Sara, et autres
Publié: (2025)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
par: Yang, Chenchen, et autres
Publié: (2026)
par: Yang, Chenchen, et autres
Publié: (2026)
Efficient Streaming LLM for Speech Recognition
par: Jia, Junteng, et autres
Publié: (2024)
par: Jia, Junteng, et autres
Publié: (2024)
SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
par: Gong, Hongyu, et autres
Publié: (2024)
par: Gong, Hongyu, et autres
Publié: (2024)
SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models
par: Yao, Wenhan, et autres
Publié: (2025)
par: Yao, Wenhan, et autres
Publié: (2025)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
par: Hu, Yuchen, et autres
Publié: (2024)
par: Hu, Yuchen, et autres
Publié: (2024)
Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
par: Liu, Tianchi, et autres
Publié: (2024)
par: Liu, Tianchi, et autres
Publié: (2024)
BayLing 2: A Multilingual Large Language Model with Efficient Language Alignment
par: Zhang, Shaolei, et autres
Publié: (2024)
par: Zhang, Shaolei, et autres
Publié: (2024)
FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
par: Yao, Yiqun, et autres
Publié: (2025)
par: Yao, Yiqun, et autres
Publié: (2025)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
par: Hu, Shujie, et autres
Publié: (2024)
par: Hu, Shujie, et autres
Publié: (2024)
Documents similaires
-
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
par: Zhang, Shaolei, et autres
Publié: (2025) -
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
par: Fang, Qingkai, et autres
Publié: (2025) -
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
par: Fang, Qingkai, et autres
Publié: (2024) -
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
par: Fang, Qingkai, et autres
Publié: (2024) -
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
par: Zhang, Shaolei, et autres
Publié: (2024)