Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Minsu, Choi, Jeongsoo, Kim, Dahun, Ro, Yong Man |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific Knowledge
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
Textless Speech-to-Speech Translation With Limited Parallel Data
von: Diwan, Anuj, et al.
Veröffentlicht: (2023)
von: Diwan, Anuj, et al.
Veröffentlicht: (2023)
AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023)
Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2024)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2024)
Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
Prompt Tuning of Deep Neural Networks for Speaker-adaptive Visual Speech Recognition
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
Compact Speech Translation Models via Discrete Speech Units Pretraining
von: Lam, Tsz Kin, et al.
Veröffentlicht: (2024)
von: Lam, Tsz Kin, et al.
Veröffentlicht: (2024)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
MultiGen: Child-Friendly Multilingual Speech Generator with LLMs
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2025)
Visual-Aware Speech Recognition for Noisy Scenarios
von: Balaji, Lakshmipathi, et al.
Veröffentlicht: (2025)
von: Balaji, Lakshmipathi, et al.
Veröffentlicht: (2025)
A Large-Scale Evaluation of Speech Foundation Models
von: Yang, Shu-wen, et al.
Veröffentlicht: (2024)
von: Yang, Shu-wen, et al.
Veröffentlicht: (2024)
Estimating the Completeness of Discrete Speech Units
von: Yeh, Sung-Lin, et al.
Veröffentlicht: (2024)
von: Yeh, Sung-Lin, et al.
Veröffentlicht: (2024)
Improved Cross-Lingual Transfer Learning For Automatic Speech Translation
von: Khurana, Sameer, et al.
Veröffentlicht: (2023)
von: Khurana, Sameer, et al.
Veröffentlicht: (2023)
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
von: Wills, Simone, et al.
Veröffentlicht: (2023)
von: Wills, Simone, et al.
Veröffentlicht: (2023)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
SpeechMLC: Speech Multi-label Classification
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation
von: Goswami, Mandip
Veröffentlicht: (2026)
von: Goswami, Mandip
Veröffentlicht: (2026)
WST-X Series: Wavelet Scattering Transform for Interpretable Speech Deepfake Detection
von: Xuan, Xi, et al.
Veröffentlicht: (2026)
von: Xuan, Xi, et al.
Veröffentlicht: (2026)
WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection
von: Xuan, Xi, et al.
Veröffentlicht: (2025)
von: Xuan, Xi, et al.
Veröffentlicht: (2025)
Text-driven Talking Face Synthesis by Reprogramming Audio-driven Models
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2024)
von: Wang, Peidong, et al.
Veröffentlicht: (2024)
Exploring Phonetic Context-Aware Lip-Sync For Talking Face Generation
von: Park, Se Jin, et al.
Veröffentlicht: (2023)
von: Park, Se Jin, et al.
Veröffentlicht: (2023)
Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2023)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2023)
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
von: Le-Duc, Khai, et al.
Veröffentlicht: (2025)
von: Le-Duc, Khai, et al.
Veröffentlicht: (2025)
Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation
von: Kim, Miseul, et al.
Veröffentlicht: (2024)
von: Kim, Miseul, et al.
Veröffentlicht: (2024)
One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech
von: Abebe, Amanuel Gizachew, et al.
Veröffentlicht: (2026)
von: Abebe, Amanuel Gizachew, et al.
Veröffentlicht: (2026)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
Speech Enhancement based on cascaded two flows
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
von: Wang, Yuejiao, et al.
Veröffentlicht: (2024)
von: Wang, Yuejiao, et al.
Veröffentlicht: (2024)
Crossmodal ASR Error Correction with Discrete Speech Units
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
von: Labrak, Yanis, et al.
Veröffentlicht: (2025)
von: Labrak, Yanis, et al.
Veröffentlicht: (2025)
DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
von: Lee, Dongheon, et al.
Veröffentlicht: (2023)
von: Lee, Dongheon, et al.
Veröffentlicht: (2023)
TMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different Languages
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
Crowdsourced Multilingual Speech Intelligibility Testing
von: Lechler, Laura, et al.
Veröffentlicht: (2024)
von: Lechler, Laura, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025) -
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
von: Duret, Jarod, et al.
Veröffentlicht: (2024) -
Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific Knowledge
von: Kim, Minsu, et al.
Veröffentlicht: (2023) -
Textless Speech-to-Speech Translation With Limited Parallel Data
von: Diwan, Anuj, et al.
Veröffentlicht: (2023) -
AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)