HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Nishimura, Yuto, Hirose, Takumi, Ohi, Masanari, Nakayama, Hideki, Inoue, Nakamasa |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
di: Inoue, Nakamasa, et al.
Pubblicazione: (2024)
di: Inoue, Nakamasa, et al.
Pubblicazione: (2024)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
di: Chen, Sanyuan, et al.
Pubblicazione: (2024)
di: Chen, Sanyuan, et al.
Pubblicazione: (2024)
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
di: Guo, Haohan, et al.
Pubblicazione: (2024)
di: Guo, Haohan, et al.
Pubblicazione: (2024)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
di: Xue, Jinlong, et al.
Pubblicazione: (2024)
di: Xue, Jinlong, et al.
Pubblicazione: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
di: Jiang, Yuepeng, et al.
Pubblicazione: (2024)
di: Jiang, Yuepeng, et al.
Pubblicazione: (2024)
A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation
di: Pei, Hanchen, et al.
Pubblicazione: (2026)
di: Pei, Hanchen, et al.
Pubblicazione: (2026)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2025)
di: Inoue, Sho, et al.
Pubblicazione: (2025)
Personalized Neural Speech Codec
di: Jang, Inseon, et al.
Pubblicazione: (2024)
di: Jang, Inseon, et al.
Pubblicazione: (2024)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
di: Lei, Shun, et al.
Pubblicazione: (2023)
di: Lei, Shun, et al.
Pubblicazione: (2023)
Zero-Shot Text-to-Speech from Continuous Text Streams
di: Dang, Trung, et al.
Pubblicazione: (2024)
di: Dang, Trung, et al.
Pubblicazione: (2024)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
di: Xin, Detai, et al.
Pubblicazione: (2024)
di: Xin, Detai, et al.
Pubblicazione: (2024)
Hierarchical Control of Emotion Rendering in Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
Probing the Robustness Properties of Neural Speech Codecs
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2025)
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2025)
SpatialCodec: Neural Spatial Speech Coding
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2023)
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2023)
A Neural Speech Codec for Noise Robust Speech Coding
di: Huang, Jiayi, et al.
Pubblicazione: (2023)
di: Huang, Jiayi, et al.
Pubblicazione: (2023)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
di: Wang, Chunhui, et al.
Pubblicazione: (2024)
di: Wang, Chunhui, et al.
Pubblicazione: (2024)
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
di: Wang, Yuancheng, et al.
Pubblicazione: (2024)
di: Wang, Yuancheng, et al.
Pubblicazione: (2024)
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
di: Wang, Hankun, et al.
Pubblicazione: (2025)
di: Wang, Hankun, et al.
Pubblicazione: (2025)
CodecFake+: A Large-Scale Neural Audio Codec-Based Deepfake Speech Dataset
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
di: Nespoli, Francesco, et al.
Pubblicazione: (2024)
di: Nespoli, Francesco, et al.
Pubblicazione: (2024)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
di: Zhang, Leying, et al.
Pubblicazione: (2025)
di: Zhang, Leying, et al.
Pubblicazione: (2025)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
di: Li, Jiaqi, et al.
Pubblicazione: (2025)
di: Li, Jiaqi, et al.
Pubblicazione: (2025)
Fewer-token Neural Speech Codec with Time-invariant Codes
di: Ren, Yong, et al.
Pubblicazione: (2023)
di: Ren, Yong, et al.
Pubblicazione: (2023)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
Zero-Shot Text-to-Speech for Vietnamese
di: Vu, Thi, et al.
Pubblicazione: (2025)
di: Vu, Thi, et al.
Pubblicazione: (2025)
LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
di: Zhao, Xiaohan, et al.
Pubblicazione: (2025)
di: Zhao, Xiaohan, et al.
Pubblicazione: (2025)
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
di: Zhou, Shuoyi, et al.
Pubblicazione: (2024)
di: Zhou, Shuoyi, et al.
Pubblicazione: (2024)
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2023)
di: Ji, Shengpeng, et al.
Pubblicazione: (2023)
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
di: Jiang, Ziyue, et al.
Pubblicazione: (2023)
di: Jiang, Ziyue, et al.
Pubblicazione: (2023)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
di: Li, Yinghao Aaron, et al.
Pubblicazione: (2024)
di: Li, Yinghao Aaron, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
di: Inoue, Nakamasa, et al.
Pubblicazione: (2024) -
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024) -
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024) -
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
di: Chen, Sanyuan, et al.
Pubblicazione: (2024) -
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
di: Guo, Haohan, et al.
Pubblicazione: (2024)