Audiobook-CC: Controllable Long-context Speech Generation for Multicast Audiobook
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Min, Yin, JingJing, Zhang, Xiang, Hao, Siyu, Hu, Yanni, Lin, Bin, Feng, Yuan, Zhou, Hongbin, Ye, Jianhao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Prosody Analysis of Audiobooks
di: Pethe, Charuta, et al.
Pubblicazione: (2023)
di: Pethe, Charuta, et al.
Pubblicazione: (2023)
Text-aware and Context-aware Expressive Audiobook Speech Synthesis
di: Guo, Dake, et al.
Pubblicazione: (2024)
di: Guo, Dake, et al.
Pubblicazione: (2024)
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
di: Rong, Yan, et al.
Pubblicazione: (2025)
di: Rong, Yan, et al.
Pubblicazione: (2025)
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
di: Park, Kyeongman, et al.
Pubblicazione: (2025)
di: Park, Kyeongman, et al.
Pubblicazione: (2025)
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
di: Pan, Yu, et al.
Pubblicazione: (2025)
di: Pan, Yu, et al.
Pubblicazione: (2025)
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
di: Yao, Jixun, et al.
Pubblicazione: (2025)
di: Yao, Jixun, et al.
Pubblicazione: (2025)
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
di: Dai, Ziqi, et al.
Pubblicazione: (2025)
di: Dai, Ziqi, et al.
Pubblicazione: (2025)
Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
di: Chen, Sijing, et al.
Pubblicazione: (2024)
di: Chen, Sijing, et al.
Pubblicazione: (2024)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
di: Yen, Hao, et al.
Pubblicazione: (2024)
di: Yen, Hao, et al.
Pubblicazione: (2024)
A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration
di: Selvamani, Shaja Arul, et al.
Pubblicazione: (2025)
di: Selvamani, Shaja Arul, et al.
Pubblicazione: (2025)
UniTalker: Conversational Speech-Visual Synthesis
di: Hu, Yifan, et al.
Pubblicazione: (2025)
di: Hu, Yifan, et al.
Pubblicazione: (2025)
VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec
di: Yang, Leyan, et al.
Pubblicazione: (2026)
di: Yang, Leyan, et al.
Pubblicazione: (2026)
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
di: Li, Ruiqi, et al.
Pubblicazione: (2026)
di: Li, Ruiqi, et al.
Pubblicazione: (2026)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
di: Tian, Jingguang, et al.
Pubblicazione: (2024)
di: Tian, Jingguang, et al.
Pubblicazione: (2024)
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
di: Pan, Yu, et al.
Pubblicazione: (2023)
di: Pan, Yu, et al.
Pubblicazione: (2023)
Magnetoencephalography (MEG) Based Non-Invasive Chinese Speech Decoding
di: Jia, Zhihong, et al.
Pubblicazione: (2025)
di: Jia, Zhihong, et al.
Pubblicazione: (2025)
PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
di: Rong, Xiaobin, et al.
Pubblicazione: (2025)
di: Rong, Xiaobin, et al.
Pubblicazione: (2025)
FNSE-SBGAN: Far-field Speech Enhancement with Schrodinger Bridge and Generative Adversarial Networks
di: Lei, Tong, et al.
Pubblicazione: (2025)
di: Lei, Tong, et al.
Pubblicazione: (2025)
Decoding Speech Envelopes from Electroencephalogram with a Contrastive Pearson Correlation Coefficient Loss
di: Liang, Yayun, et al.
Pubblicazione: (2026)
di: Liang, Yayun, et al.
Pubblicazione: (2026)
Clever Hans Effect Found in Automatic Detection of Alzheimer's Disease through Speech
di: Liu, Yin-Long, et al.
Pubblicazione: (2024)
di: Liu, Yin-Long, et al.
Pubblicazione: (2024)
Rethinking Flow and Diffusion Bridge Models for Speech Enhancement
di: Wang, Dahan, et al.
Pubblicazione: (2026)
di: Wang, Dahan, et al.
Pubblicazione: (2026)
End-to-End DOA-Guided Speech Extraction in Noisy Multi-Talker Scenarios
di: Jing, Kangqi, et al.
Pubblicazione: (2025)
di: Jing, Kangqi, et al.
Pubblicazione: (2025)
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
di: Hu, Yifan, et al.
Pubblicazione: (2025)
di: Hu, Yifan, et al.
Pubblicazione: (2025)
Best Audiobooks of 2006
di: Burns, Ann
Pubblicazione: (2007)
di: Burns, Ann
Pubblicazione: (2007)
Best Audiobooks of 2008
di: Kuzyk, Raya
Pubblicazione: (2009)
di: Kuzyk, Raya
Pubblicazione: (2009)
A Lightweight Hybrid Dual Channel Speech Enhancement System under Low-SNR Conditions
di: Wang, Zheng, et al.
Pubblicazione: (2025)
di: Wang, Zheng, et al.
Pubblicazione: (2025)
Towards Data Drift Monitoring for Speech Deepfake Detection in the context of MLOps
di: Wang, Xin, et al.
Pubblicazione: (2025)
di: Wang, Xin, et al.
Pubblicazione: (2025)
Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild
di: Tzeng, Jing-Tong, et al.
Pubblicazione: (2025)
di: Tzeng, Jing-Tong, et al.
Pubblicazione: (2025)
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
di: Jing, Xin, et al.
Pubblicazione: (2024)
di: Jing, Xin, et al.
Pubblicazione: (2024)
Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025)
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025)
Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models
di: Tao, Dehua, et al.
Pubblicazione: (2026)
di: Tao, Dehua, et al.
Pubblicazione: (2026)
Adaptive Convolution for CNN-based Speech Enhancement Models
di: Wang, Dahan, et al.
Pubblicazione: (2025)
di: Wang, Dahan, et al.
Pubblicazione: (2025)
Top Audiobooks of 2009
di: Kuzyk, Raya
Pubblicazione: (2010)
di: Kuzyk, Raya
Pubblicazione: (2010)
LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
di: Zhao, Xiaohan, et al.
Pubblicazione: (2025)
di: Zhao, Xiaohan, et al.
Pubblicazione: (2025)
UL-UNAS: Ultra-Lightweight U-Nets for Real-Time Speech Enhancement via Network Architecture Search
di: Rong, Xiaobin, et al.
Pubblicazione: (2025)
di: Rong, Xiaobin, et al.
Pubblicazione: (2025)
Leveraging Local and Global Knowledge Integration with Time-Frequency Calibrated Distillation for Speech Enhancement
di: Cheng, Jiaming, et al.
Pubblicazione: (2025)
di: Cheng, Jiaming, et al.
Pubblicazione: (2025)
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
di: Wan, Genshun, et al.
Pubblicazione: (2026)
di: Wan, Genshun, et al.
Pubblicazione: (2026)
SNR-Progressive Model with Harmonic Compensation for Low-SNR Speech Enhancement
di: Hou, Zhongshu, et al.
Pubblicazione: (2024)
di: Hou, Zhongshu, et al.
Pubblicazione: (2024)
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
di: Li, Hanzhao, et al.
Pubblicazione: (2024)
di: Li, Hanzhao, et al.
Pubblicazione: (2024)
Beyond Manual Transcripts: The Potential of Automated Speech Recognition Errors in Improving Alzheimer's Disease Detection
di: Liu, Yin-Long, et al.
Pubblicazione: (2025)
di: Liu, Yin-Long, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Prosody Analysis of Audiobooks
di: Pethe, Charuta, et al.
Pubblicazione: (2023) -
Text-aware and Context-aware Expressive Audiobook Speech Synthesis
di: Guo, Dake, et al.
Pubblicazione: (2024) -
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
di: Rong, Yan, et al.
Pubblicazione: (2025) -
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
di: Park, Kyeongman, et al.
Pubblicazione: (2025) -
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
di: Pan, Yu, et al.
Pubblicazione: (2025)