CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Xiaosu, Sun, Zihan, Jia, Peilei, Gao, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Utterance Conditioned VAE for Speech Generation
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
by: Liu, Zhanxun, et al.
Published: (2025)
by: Liu, Zhanxun, et al.
Published: (2025)
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
by: Zhao, Yiyang, et al.
Published: (2025)
by: Zhao, Yiyang, et al.
Published: (2025)
Hello-Chat: Towards Realistic Social Audio Interactions
by: Hou, Yueran, et al.
Published: (2026)
by: Hou, Yueran, et al.
Published: (2026)
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
by: Zhang, Xueyao, et al.
Published: (2025)
by: Zhang, Xueyao, et al.
Published: (2025)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
by: Zheng, Zhisheng, et al.
Published: (2025)
by: Zheng, Zhisheng, et al.
Published: (2025)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
by: Cheng, Sitong, et al.
Published: (2025)
by: Cheng, Sitong, et al.
Published: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
by: Anastassiou, Philip, et al.
Published: (2024)
by: Anastassiou, Philip, et al.
Published: (2024)
CapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation
by: Chu, Xuangeng, et al.
Published: (2026)
by: Chu, Xuangeng, et al.
Published: (2026)
TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis
by: Liang, Qifan, et al.
Published: (2026)
by: Liang, Qifan, et al.
Published: (2026)
Multi-Utterance Speech Separation and Association Trained on Short Segments
by: Wang, Yuzhu, et al.
Published: (2025)
by: Wang, Yuzhu, et al.
Published: (2025)
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
by: Flynn, Robert, et al.
Published: (2026)
by: Flynn, Robert, et al.
Published: (2026)
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
by: Wang, Yuzhu, et al.
Published: (2025)
by: Wang, Yuzhu, et al.
Published: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
by: Li, Hebeizi, et al.
Published: (2026)
by: Li, Hebeizi, et al.
Published: (2026)
On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation
by: Cheng, Changhao, et al.
Published: (2026)
by: Cheng, Changhao, et al.
Published: (2026)
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
by: Xue, Jun, et al.
Published: (2026)
by: Xue, Jun, et al.
Published: (2026)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
by: Huo, Mingyue, et al.
Published: (2025)
by: Huo, Mingyue, et al.
Published: (2025)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
by: Zhang, Enshi, et al.
Published: (2024)
by: Zhang, Enshi, et al.
Published: (2024)
VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis
by: Ma, Chengyuan, et al.
Published: (2026)
by: Ma, Chengyuan, et al.
Published: (2026)
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
by: Zhou, Yixuan, et al.
Published: (2025)
by: Zhou, Yixuan, et al.
Published: (2025)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
by: Zink, Oswald, et al.
Published: (2024)
by: Zink, Oswald, et al.
Published: (2024)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
by: Huang, Kuan-Po, et al.
Published: (2023)
by: Huang, Kuan-Po, et al.
Published: (2023)
Adapting Speech Language Model to Singing Voice Synthesis
by: Zhao, Yiwen, et al.
Published: (2025)
by: Zhao, Yiwen, et al.
Published: (2025)
Perpetual Dialogues: A Computational Analysis of Voice-Guitar Interaction in Carlos Paredes's Discography
by: Bernardes, Gilberto, et al.
Published: (2026)
by: Bernardes, Gilberto, et al.
Published: (2026)
VoiceBridge: General Speech Restoration with One-step Latent Bridge Models
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
by: Yang, Runyan, et al.
Published: (2024)
by: Yang, Runyan, et al.
Published: (2024)
Cross-Talk Speech Reduction, by Separation, for Separation
by: Wang, Zhong-Qiu, et al.
Published: (2026)
by: Wang, Zhong-Qiu, et al.
Published: (2026)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
by: Xie, Hanke, et al.
Published: (2025)
by: Xie, Hanke, et al.
Published: (2025)
DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions
by: Niu, Shu-Tong, et al.
Published: (2024)
by: Niu, Shu-Tong, et al.
Published: (2024)
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
by: Li, Fengjin, et al.
Published: (2025)
by: Li, Fengjin, et al.
Published: (2025)
VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
by: Zhou, Yixuan, et al.
Published: (2024)
by: Zhou, Yixuan, et al.
Published: (2024)
Speech to Speech Synthesis for Voice Impersonation
by: Johnson, Bjorn, et al.
Published: (2026)
by: Johnson, Bjorn, et al.
Published: (2026)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
by: Wang, Helin, et al.
Published: (2025)
by: Wang, Helin, et al.
Published: (2025)
DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
by: Xu, Haiying, et al.
Published: (2025)
by: Xu, Haiying, et al.
Published: (2025)
Improving Short Utterance Anti-Spoofing with AASIST2
by: Zhang, Yuxiang, et al.
Published: (2023)
by: Zhang, Yuxiang, et al.
Published: (2023)
Geneses: Unified Generative Speech Enhancement and Separation
by: Asai, Kohei, et al.
Published: (2026)
by: Asai, Kohei, et al.
Published: (2026)
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
by: Kang, Fang, et al.
Published: (2025)
by: Kang, Fang, et al.
Published: (2025)
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
by: Guragain, Anmol, et al.
Published: (2024)
by: Guragain, Anmol, et al.
Published: (2024)
Similar Items
-
Cross-Utterance Conditioned VAE for Speech Generation
by: Li, Yang, et al.
Published: (2023) -
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
by: Liu, Zhanxun, et al.
Published: (2025) -
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
by: Zhao, Yiyang, et al.
Published: (2025) -
Hello-Chat: Towards Realistic Social Audio Interactions
by: Hou, Yueran, et al.
Published: (2026) -
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
by: Zhang, Xueyao, et al.
Published: (2025)