Acoustic BPE for Speech Generation with Discrete Tokens
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Shen, Feiyu, Guo, Yiwei, Du, Chenpeng, Chen, Xie, Yu, Kai |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
par: Du, Chenpeng, et autres
Publié: (2022)
par: Du, Chenpeng, et autres
Publié: (2022)
On the Effectiveness of Acoustic BPE in Decoder-Only TTS
par: Li, Bohan, et autres
Publié: (2024)
par: Li, Bohan, et autres
Publié: (2024)
Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge
par: Du, Chenpeng, et autres
Publié: (2023)
par: Du, Chenpeng, et autres
Publié: (2023)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
par: Guo, Yiwei, et autres
Publié: (2024)
par: Guo, Yiwei, et autres
Publié: (2024)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
par: Du, Chenpeng, et autres
Publié: (2023)
par: Du, Chenpeng, et autres
Publié: (2023)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
par: Wang, Hankun, et autres
Publié: (2024)
par: Wang, Hankun, et autres
Publié: (2024)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
par: Du, Chenpeng, et autres
Publié: (2021)
par: Du, Chenpeng, et autres
Publié: (2021)
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
par: Wang, Haoran, et autres
Publié: (2025)
par: Wang, Haoran, et autres
Publié: (2025)
Recent Advances in Discrete Speech Tokens: A Review
par: Guo, Yiwei, et autres
Publié: (2025)
par: Guo, Yiwei, et autres
Publié: (2025)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
par: Du, Chenpeng, et autres
Publié: (2024)
par: Du, Chenpeng, et autres
Publié: (2024)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
par: Guo, Yiwei, et autres
Publié: (2024)
par: Guo, Yiwei, et autres
Publié: (2024)
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
par: Wang, Hankun, et autres
Publié: (2024)
par: Wang, Hankun, et autres
Publié: (2024)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
par: Guo, Yiwei, et autres
Publié: (2023)
par: Guo, Yiwei, et autres
Publié: (2023)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
par: Wang, Huimeng, et autres
Publié: (2025)
par: Wang, Huimeng, et autres
Publié: (2025)
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
par: Xie, Kun, et autres
Publié: (2025)
par: Xie, Kun, et autres
Publié: (2025)
Confidence-based Filtering for Speech Dataset Curation with Generative Speech Enhancement Using Discrete Tokens
par: Yamauchi, Kazuki, et autres
Publié: (2026)
par: Yamauchi, Kazuki, et autres
Publié: (2026)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
par: Lu, Ye-Xin, et autres
Publié: (2025)
par: Lu, Ye-Xin, et autres
Publié: (2025)
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
par: Liu, Sen, et autres
Publié: (2024)
par: Liu, Sen, et autres
Publié: (2024)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
par: Guo, Hao-Han, et autres
Publié: (2024)
par: Guo, Hao-Han, et autres
Publié: (2024)
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
par: Wang, Hankun, et autres
Publié: (2025)
par: Wang, Hankun, et autres
Publié: (2025)
HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
par: Li, Bohan, et autres
Publié: (2026)
par: Li, Bohan, et autres
Publié: (2026)
Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model
par: Wang, Siyang, et autres
Publié: (2024)
par: Wang, Siyang, et autres
Publié: (2024)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
par: Wang, Chunhui, et autres
Publié: (2024)
par: Wang, Chunhui, et autres
Publié: (2024)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
par: Chen, Peikun, et autres
Publié: (2024)
par: Chen, Peikun, et autres
Publié: (2024)
Rethinking Discrete Speech Representation Tokens for Accent Generation
par: Zhong, Jinzuomu, et autres
Publié: (2026)
par: Zhong, Jinzuomu, et autres
Publié: (2026)
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
par: Zhang, Hanglei, et autres
Publié: (2025)
par: Zhang, Hanglei, et autres
Publié: (2025)
Exploring the Benefits of Tokenization of Discrete Acoustic Units
par: Dekel, Avihu, et autres
Publié: (2024)
par: Dekel, Avihu, et autres
Publié: (2024)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
par: Tao, Dehua, et autres
Publié: (2024)
par: Tao, Dehua, et autres
Publié: (2024)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
par: Chen, Wenxi, et autres
Publié: (2025)
par: Chen, Wenxi, et autres
Publié: (2025)
From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
par: Mu, Zhaoxi, et autres
Publié: (2025)
par: Mu, Zhaoxi, et autres
Publié: (2025)
Autoregressive Speech Enhancement via Acoustic Tokens
par: Della Libera, Luca, et autres
Publié: (2025)
par: Della Libera, Luca, et autres
Publié: (2025)
DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
par: Li, Baihan, et autres
Publié: (2024)
par: Li, Baihan, et autres
Publié: (2024)
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
par: Guo, Dake, et autres
Publié: (2025)
par: Guo, Dake, et autres
Publié: (2025)
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
par: Onda, Kentaro, et autres
Publié: (2025)
par: Onda, Kentaro, et autres
Publié: (2025)
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
par: Onda, Kentaro, et autres
Publié: (2025)
par: Onda, Kentaro, et autres
Publié: (2025)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
par: Wang, Yuanyuan, et autres
Publié: (2025)
par: Wang, Yuanyuan, et autres
Publié: (2025)
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
par: Chang, Xuankai, et autres
Publié: (2024)
par: Chang, Xuankai, et autres
Publié: (2024)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
par: Guo, Hao-Han, et autres
Publié: (2025)
par: Guo, Hao-Han, et autres
Publié: (2025)
Continuous Speech Tokenizer in Text To Speech
par: Li, Yixing, et autres
Publié: (2024)
par: Li, Yixing, et autres
Publié: (2024)
Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
par: Guo, Haohan, et autres
Publié: (2024)
par: Guo, Haohan, et autres
Publié: (2024)
Documents similaires
-
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
par: Du, Chenpeng, et autres
Publié: (2022) -
On the Effectiveness of Acoustic BPE in Decoder-Only TTS
par: Li, Bohan, et autres
Publié: (2024) -
Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge
par: Du, Chenpeng, et autres
Publié: (2023) -
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
par: Guo, Yiwei, et autres
Publié: (2024) -
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
par: Du, Chenpeng, et autres
Publié: (2023)