CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Su, Xiaosu, Sun, Zihan, Jia, Peilei, Gao, Jun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Cross-Utterance Conditioned VAE for Speech Generation
di: Li, Yang, et al.
Pubblicazione: (2023)
di: Li, Yang, et al.
Pubblicazione: (2023)
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
di: Liu, Zhanxun, et al.
Pubblicazione: (2025)
di: Liu, Zhanxun, et al.
Pubblicazione: (2025)
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
di: Zhao, Yiyang, et al.
Pubblicazione: (2025)
di: Zhao, Yiyang, et al.
Pubblicazione: (2025)
Hello-Chat: Towards Realistic Social Audio Interactions
di: Hou, Yueran, et al.
Pubblicazione: (2026)
di: Hou, Yueran, et al.
Pubblicazione: (2026)
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
di: Zheng, Zhisheng, et al.
Pubblicazione: (2025)
di: Zheng, Zhisheng, et al.
Pubblicazione: (2025)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
di: Cheng, Sitong, et al.
Pubblicazione: (2025)
di: Cheng, Sitong, et al.
Pubblicazione: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
CapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation
di: Chu, Xuangeng, et al.
Pubblicazione: (2026)
di: Chu, Xuangeng, et al.
Pubblicazione: (2026)
TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis
di: Liang, Qifan, et al.
Pubblicazione: (2026)
di: Liang, Qifan, et al.
Pubblicazione: (2026)
Multi-Utterance Speech Separation and Association Trained on Short Segments
di: Wang, Yuzhu, et al.
Pubblicazione: (2025)
di: Wang, Yuzhu, et al.
Pubblicazione: (2025)
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
di: Flynn, Robert, et al.
Pubblicazione: (2026)
di: Flynn, Robert, et al.
Pubblicazione: (2026)
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
di: Wang, Yuzhu, et al.
Pubblicazione: (2025)
di: Wang, Yuzhu, et al.
Pubblicazione: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
di: Li, Hebeizi, et al.
Pubblicazione: (2026)
di: Li, Hebeizi, et al.
Pubblicazione: (2026)
On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation
di: Cheng, Changhao, et al.
Pubblicazione: (2026)
di: Cheng, Changhao, et al.
Pubblicazione: (2026)
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
di: Xue, Jun, et al.
Pubblicazione: (2026)
di: Xue, Jun, et al.
Pubblicazione: (2026)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
di: Zhang, Enshi, et al.
Pubblicazione: (2024)
di: Zhang, Enshi, et al.
Pubblicazione: (2024)
VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis
di: Ma, Chengyuan, et al.
Pubblicazione: (2026)
di: Ma, Chengyuan, et al.
Pubblicazione: (2026)
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
di: Zhou, Yixuan, et al.
Pubblicazione: (2025)
di: Zhou, Yixuan, et al.
Pubblicazione: (2025)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
di: Zink, Oswald, et al.
Pubblicazione: (2024)
di: Zink, Oswald, et al.
Pubblicazione: (2024)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
di: Huang, Kuan-Po, et al.
Pubblicazione: (2023)
di: Huang, Kuan-Po, et al.
Pubblicazione: (2023)
Adapting Speech Language Model to Singing Voice Synthesis
di: Zhao, Yiwen, et al.
Pubblicazione: (2025)
di: Zhao, Yiwen, et al.
Pubblicazione: (2025)
Perpetual Dialogues: A Computational Analysis of Voice-Guitar Interaction in Carlos Paredes's Discography
di: Bernardes, Gilberto, et al.
Pubblicazione: (2026)
di: Bernardes, Gilberto, et al.
Pubblicazione: (2026)
VoiceBridge: General Speech Restoration with One-step Latent Bridge Models
di: Zhang, Chi, et al.
Pubblicazione: (2025)
di: Zhang, Chi, et al.
Pubblicazione: (2025)
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
di: Yang, Runyan, et al.
Pubblicazione: (2024)
di: Yang, Runyan, et al.
Pubblicazione: (2024)
Cross-Talk Speech Reduction, by Separation, for Separation
di: Wang, Zhong-Qiu, et al.
Pubblicazione: (2026)
di: Wang, Zhong-Qiu, et al.
Pubblicazione: (2026)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
di: Xie, Hanke, et al.
Pubblicazione: (2025)
di: Xie, Hanke, et al.
Pubblicazione: (2025)
DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions
di: Niu, Shu-Tong, et al.
Pubblicazione: (2024)
di: Niu, Shu-Tong, et al.
Pubblicazione: (2024)
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
di: Li, Fengjin, et al.
Pubblicazione: (2025)
di: Li, Fengjin, et al.
Pubblicazione: (2025)
VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
di: Zhou, Yixuan, et al.
Pubblicazione: (2024)
di: Zhou, Yixuan, et al.
Pubblicazione: (2024)
Speech to Speech Synthesis for Voice Impersonation
di: Johnson, Bjorn, et al.
Pubblicazione: (2026)
di: Johnson, Bjorn, et al.
Pubblicazione: (2026)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
di: Wang, Helin, et al.
Pubblicazione: (2025)
di: Wang, Helin, et al.
Pubblicazione: (2025)
DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue
di: Li, Xiang, et al.
Pubblicazione: (2025)
di: Li, Xiang, et al.
Pubblicazione: (2025)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
di: Shi, Weiyan, et al.
Pubblicazione: (2025)
di: Shi, Weiyan, et al.
Pubblicazione: (2025)
EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
di: Xu, Haiying, et al.
Pubblicazione: (2025)
di: Xu, Haiying, et al.
Pubblicazione: (2025)
Improving Short Utterance Anti-Spoofing with AASIST2
di: Zhang, Yuxiang, et al.
Pubblicazione: (2023)
di: Zhang, Yuxiang, et al.
Pubblicazione: (2023)
Geneses: Unified Generative Speech Enhancement and Separation
di: Asai, Kohei, et al.
Pubblicazione: (2026)
di: Asai, Kohei, et al.
Pubblicazione: (2026)
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
di: Kang, Fang, et al.
Pubblicazione: (2025)
di: Kang, Fang, et al.
Pubblicazione: (2025)
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
di: Guragain, Anmol, et al.
Pubblicazione: (2024)
di: Guragain, Anmol, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Cross-Utterance Conditioned VAE for Speech Generation
di: Li, Yang, et al.
Pubblicazione: (2023) -
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
di: Liu, Zhanxun, et al.
Pubblicazione: (2025) -
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
di: Zhao, Yiyang, et al.
Pubblicazione: (2025) -
Hello-Chat: Towards Realistic Social Audio Interactions
di: Hou, Yueran, et al.
Pubblicazione: (2026) -
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)