SLM-S2ST: A multimodal language model for direct speech-to-speech translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Yuxuan, Wu, Haibin, Fan, Ruchao, Wang, Xiaofei, Lu, Heng, Qian, Yao, Li, Jinyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
von: Zhao, Rui, et al.
Veröffentlicht: (2024)
von: Zhao, Rui, et al.
Veröffentlicht: (2024)
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
von: Wu, Haibin, et al.
Veröffentlicht: (2025)
von: Wu, Haibin, et al.
Veröffentlicht: (2025)
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
Omni-directional attention mechanism based on Mamba for speech separation
von: Xue, Ke, et al.
Veröffentlicht: (2026)
von: Xue, Ke, et al.
Veröffentlicht: (2026)
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models
von: Ren, Bo, et al.
Veröffentlicht: (2026)
von: Ren, Bo, et al.
Veröffentlicht: (2026)
Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio
von: Shi, Mohan, et al.
Veröffentlicht: (2025)
von: Shi, Mohan, et al.
Veröffentlicht: (2025)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
DBMIF: a deep balanced multimodal iterative fusion framework for air- and bone-conduction speech enhancement
von: Wu, Yilei, et al.
Veröffentlicht: (2026)
von: Wu, Yilei, et al.
Veröffentlicht: (2026)
Transcribe, Align and Segment: Creating speech datasets for low-resource languages
von: Sereda, Taras
Veröffentlicht: (2024)
von: Sereda, Taras
Veröffentlicht: (2024)
Transferable speech-to-text large language model alignment module
von: Wu, Boyong, et al.
Veröffentlicht: (2024)
von: Wu, Boyong, et al.
Veröffentlicht: (2024)
Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
von: Shirahata, Yuma, et al.
Veröffentlicht: (2024)
von: Shirahata, Yuma, et al.
Veröffentlicht: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
von: Saon, George, et al.
Veröffentlicht: (2025)
von: Saon, George, et al.
Veröffentlicht: (2025)
Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions
von: Mack, Wolfgang, et al.
Veröffentlicht: (2025)
von: Mack, Wolfgang, et al.
Veröffentlicht: (2025)
Predicting speech intelligibility in older adults for speech enhancement using the Gammachirp Envelope Similarity Index, GESI
von: Yamamoto, Ayako, et al.
Veröffentlicht: (2025)
von: Yamamoto, Ayako, et al.
Veröffentlicht: (2025)
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
von: An, Keyu, et al.
Veröffentlicht: (2025)
von: An, Keyu, et al.
Veröffentlicht: (2025)
Good practices for evaluation of synthesized speech
von: Cooper, Erica, et al.
Veröffentlicht: (2025)
von: Cooper, Erica, et al.
Veröffentlicht: (2025)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
von: Deng, Qingkun, et al.
Veröffentlicht: (2024)
von: Deng, Qingkun, et al.
Veröffentlicht: (2024)
CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
End-to-end transfer learning for speaker-independent cross-language and cross-corpus speech emotion recognition
von: Tang, Duowei, et al.
Veröffentlicht: (2023)
von: Tang, Duowei, et al.
Veröffentlicht: (2023)
Prominence-aware automatic speech recognition for conversational speech
von: Linke, Julian, et al.
Veröffentlicht: (2025)
von: Linke, Julian, et al.
Veröffentlicht: (2025)
Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation
von: Ma, Jianbo, et al.
Veröffentlicht: (2026)
von: Ma, Jianbo, et al.
Veröffentlicht: (2026)
Exploring speech style spaces with language models: Emotional TTS without emotion labels
von: Chandra, Shreeram Suresh, et al.
Veröffentlicht: (2024)
von: Chandra, Shreeram Suresh, et al.
Veröffentlicht: (2024)
Direct Punjabi to English speech translation using discrete units
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024)
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024)
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
von: Akkiraju, Bhavana, et al.
Veröffentlicht: (2025)
von: Akkiraju, Bhavana, et al.
Veröffentlicht: (2025)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
Probing mental health information in speech foundation models
von: de Gennes, Marc, et al.
Veröffentlicht: (2024)
von: de Gennes, Marc, et al.
Veröffentlicht: (2024)
WhisperFlow: speech foundation models in real time
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024)
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024)
End-to-end multi-channel speaker extraction and binaural speech synthesis
von: Chi, Cheng, et al.
Veröffentlicht: (2024)
von: Chi, Cheng, et al.
Veröffentlicht: (2024)
On the relationship between speech and hearing
von: Umesh, Srinivasan, et al.
Veröffentlicht: (2024)
von: Umesh, Srinivasan, et al.
Veröffentlicht: (2024)
Index-MSR: A high-efficiency multimodal fusion framework for speech recognition
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
Do self-supervised speech and language models extract similar representations as human brain?
von: Chen, Peili, et al.
Veröffentlicht: (2023)
von: Chen, Peili, et al.
Veröffentlicht: (2023)
Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement
von: Lee, Dongheon, et al.
Veröffentlicht: (2026)
von: Lee, Dongheon, et al.
Veröffentlicht: (2026)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
von: Kesiraju, Santosh, et al.
Veröffentlicht: (2023)
von: Kesiraju, Santosh, et al.
Veröffentlicht: (2023)
TS3-Codec: Transformer-Based Simple Streaming Single Codec
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
On the social bias of speech self-supervised models
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
von: Gubian, Michele, et al.
Veröffentlicht: (2025)
von: Gubian, Michele, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
von: Zhao, Rui, et al.
Veröffentlicht: (2024) -
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
von: Wu, Haibin, et al.
Veröffentlicht: (2025) -
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM
von: Fan, Ruchao, et al.
Veröffentlicht: (2024) -
Omni-directional attention mechanism based on Mamba for speech separation
von: Xue, Ke, et al.
Veröffentlicht: (2026) -
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
von: Huang, Ziling, et al.
Veröffentlicht: (2025)