CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Ji-Hoon, Yang, Hong-Sun, Ju, Yoon-Cheol, Kim, Il-Hwan, Kim, Byeong-Yeol, Chung, Joon Son |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
von: Jang, Youngjoon, et al.
Veröffentlicht: (2024)
von: Jang, Youngjoon, et al.
Veröffentlicht: (2024)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2023)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2023)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
von: Kwak, Doyeop, et al.
Veröffentlicht: (2025)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2025)
Lightweight Audio Segmentation for Long-form Speech Translation
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
von: Park, Nohil, et al.
Veröffentlicht: (2024)
von: Park, Nohil, et al.
Veröffentlicht: (2024)
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2024)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2024)
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
von: Kim, Minje, et al.
Veröffentlicht: (2024)
von: Kim, Minje, et al.
Veröffentlicht: (2024)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
Boosting Unknown-number Speaker Separation with Transformer Decoder-based Attractor
von: Lee, Younglo, et al.
Veröffentlicht: (2024)
von: Lee, Younglo, et al.
Veröffentlicht: (2024)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
Semantic Communications for Speech Recognition
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
Ultrasensitive Textile Strain Sensors Redefine Wearable Silent Speech Interfaces with High Machine Learning Efficiency
von: Tang, Chenyu, et al.
Veröffentlicht: (2023)
von: Tang, Chenyu, et al.
Veröffentlicht: (2023)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
von: Bae, Hanbin, et al.
Veröffentlicht: (2024)
von: Bae, Hanbin, et al.
Veröffentlicht: (2024)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
A Study on Speech Assessment with Visual Cues
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
Toward Universal Speech Enhancement for Diverse Input Conditions
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
Speech dereverberation constrained on room impulse response characteristics
von: Bahrman, Louis, et al.
Veröffentlicht: (2024)
von: Bahrman, Louis, et al.
Veröffentlicht: (2024)
Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech
von: Maghsoudi, Maryam, et al.
Veröffentlicht: (2026)
von: Maghsoudi, Maryam, et al.
Veröffentlicht: (2026)
EEG-Based Speech Decoding: A Novel Approach Using Multi-Kernel Ensemble Diffusion Models
von: Kim, Soowon, et al.
Veröffentlicht: (2024)
von: Kim, Soowon, et al.
Veröffentlicht: (2024)
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
von: Kim, Youngjae, et al.
Veröffentlicht: (2024)
von: Kim, Youngjae, et al.
Veröffentlicht: (2024)
Significance of Chirp MFCC as a Feature in Speech and Audio Applications
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
On Improving Error Resilience of Neural End-to-End Speech Coders
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
von: Zhang, Wangyou, et al.
Veröffentlicht: (2025)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2025)
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
von: Rahimi, Akam, et al.
Veröffentlicht: (2025)
von: Rahimi, Akam, et al.
Veröffentlicht: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
von: Kwon, Younghoo, et al.
Veröffentlicht: (2024)
von: Kwon, Younghoo, et al.
Veröffentlicht: (2024)
Cross-Talk Reduction
von: Wang, Zhong-Qiu, et al.
Veröffentlicht: (2024)
von: Wang, Zhong-Qiu, et al.
Veröffentlicht: (2024)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
von: Jang, Youngjoon, et al.
Veröffentlicht: (2024) -
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024) -
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2023) -
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025) -
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)