RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
Fuente:
arXiv
Guardado en:
| Autores principales: | Xin, Detai, Tan, Xu, Shen, Kai, Ju, Zeqian, Yang, Dongchao, Wang, Yuancheng, Takamichi, Shinnosuke, Saruwatari, Hiroshi, Liu, Shujie, Li, Jinyu, Zhao, Sheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
por: Xin, Detai, et al.
Publicado: (2024)
por: Xin, Detai, et al.
Publicado: (2024)
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
por: Seki, Kentaro, et al.
Publicado: (2025)
por: Seki, Kentaro, et al.
Publicado: (2025)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
por: Xin, Detai, et al.
Publicado: (2023)
por: Xin, Detai, et al.
Publicado: (2023)
Building speech corpus with diverse voice characteristics for its prompt-based representation
por: Watanabe, Aya, et al.
Publicado: (2024)
por: Watanabe, Aya, et al.
Publicado: (2024)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
por: Saeki, Takaaki, et al.
Publicado: (2024)
por: Saeki, Takaaki, et al.
Publicado: (2024)
DNN-based ensemble singing voice synthesis with interactions between singers
por: Hyodo, Hiroaki, et al.
Publicado: (2024)
por: Hyodo, Hiroaki, et al.
Publicado: (2024)
Drum-to-Vocal Percussion Sound Conversion and Its Evaluation Methodology
por: Nobukawa, Rinka, et al.
Publicado: (2025)
por: Nobukawa, Rinka, et al.
Publicado: (2025)
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
por: Ju, Zeqian, et al.
Publicado: (2024)
por: Ju, Zeqian, et al.
Publicado: (2024)
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
por: Guo, Haohan, et al.
Publicado: (2024)
por: Guo, Haohan, et al.
Publicado: (2024)
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
por: Igarashi, Takuto, et al.
Publicado: (2024)
por: Igarashi, Takuto, et al.
Publicado: (2024)
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
por: Suda, Hitoshi, et al.
Publicado: (2025)
por: Suda, Hitoshi, et al.
Publicado: (2025)
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
por: Kanamori, Yusuke, et al.
Publicado: (2025)
por: Kanamori, Yusuke, et al.
Publicado: (2025)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
por: Seki, Kentaro, et al.
Publicado: (2024)
por: Seki, Kentaro, et al.
Publicado: (2024)
JaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
por: Nakamura, Tomohiko, et al.
Publicado: (2022)
por: Nakamura, Tomohiko, et al.
Publicado: (2022)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
por: Saito, Yuki, et al.
Publicado: (2024)
por: Saito, Yuki, et al.
Publicado: (2024)
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
por: Xin, Detai, et al.
Publicado: (2026)
por: Xin, Detai, et al.
Publicado: (2026)
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
por: Chen, Sanyuan, et al.
Publicado: (2024)
por: Chen, Sanyuan, et al.
Publicado: (2024)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
por: Nakata, Wataru, et al.
Publicado: (2024)
por: Nakata, Wataru, et al.
Publicado: (2024)
A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation
por: Pei, Hanchen, et al.
Publicado: (2026)
por: Pei, Hanchen, et al.
Publicado: (2026)
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
por: Kando, Shunsuke, et al.
Publicado: (2025)
por: Kando, Shunsuke, et al.
Publicado: (2025)
Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing
por: Nakata, Wataru, et al.
Publicado: (2025)
por: Nakata, Wataru, et al.
Publicado: (2025)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
por: Li, Jiaqi, et al.
Publicado: (2024)
por: Li, Jiaqi, et al.
Publicado: (2024)
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
por: Yang, Dong, et al.
Publicado: (2025)
por: Yang, Dong, et al.
Publicado: (2025)
Hyperbolic Embeddings for Order-Aware Classification of Audio Effect Chains
por: Wada, Aogu, et al.
Publicado: (2025)
por: Wada, Aogu, et al.
Publicado: (2025)
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
por: Wang, Yuancheng, et al.
Publicado: (2025)
por: Wang, Yuancheng, et al.
Publicado: (2025)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
por: Li, Jiaqi, et al.
Publicado: (2025)
por: Li, Jiaqi, et al.
Publicado: (2025)
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
por: Yang, Jianing, et al.
Publicado: (2025)
por: Yang, Jianing, et al.
Publicado: (2025)
Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
por: Suda, Hitoshi, et al.
Publicado: (2024)
por: Suda, Hitoshi, et al.
Publicado: (2024)
SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis
por: Take, Osamu, et al.
Publicado: (2024)
por: Take, Osamu, et al.
Publicado: (2024)
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
por: Guo, Haohan, et al.
Publicado: (2024)
por: Guo, Haohan, et al.
Publicado: (2024)
Geneses: Unified Generative Speech Enhancement and Separation
por: Asai, Kohei, et al.
Publicado: (2026)
por: Asai, Kohei, et al.
Publicado: (2026)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
por: Yamauchi, Kazuki, et al.
Publicado: (2024)
por: Yamauchi, Kazuki, et al.
Publicado: (2024)
Localizing Acoustic Energy in Sound Field Synthesis by Directionally Weighted Exterior Radiation Suppression
por: Tomita, Yoshihide, et al.
Publicado: (2024)
por: Tomita, Yoshihide, et al.
Publicado: (2024)
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
por: Han, Bing, et al.
Publicado: (2024)
por: Han, Bing, et al.
Publicado: (2024)
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
por: Kishi, Minoru, et al.
Publicado: (2025)
por: Kishi, Minoru, et al.
Publicado: (2025)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
por: Tsunoo, Emiru, et al.
Publicado: (2024)
por: Tsunoo, Emiru, et al.
Publicado: (2024)
Probing the Robustness Properties of Neural Speech Codecs
por: Tseng, Wei-Cheng, et al.
Publicado: (2025)
por: Tseng, Wei-Cheng, et al.
Publicado: (2025)
A Neural Speech Codec for Noise Robust Speech Coding
por: Huang, Jiayi, et al.
Publicado: (2023)
por: Huang, Jiayi, et al.
Publicado: (2023)
NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
por: Niu, Zhikang, et al.
Publicado: (2024)
por: Niu, Zhikang, et al.
Publicado: (2024)
Boosting Large Language Model for Speech Synthesis: An Empirical Study
por: Hao, Hongkun, et al.
Publicado: (2023)
por: Hao, Hongkun, et al.
Publicado: (2023)
Ejemplares similares
-
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
por: Xin, Detai, et al.
Publicado: (2024) -
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
por: Seki, Kentaro, et al.
Publicado: (2025) -
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
por: Xin, Detai, et al.
Publicado: (2023) -
Building speech corpus with diverse voice characteristics for its prompt-based representation
por: Watanabe, Aya, et al.
Publicado: (2024) -
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
por: Saeki, Takaaki, et al.
Publicado: (2024)