CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Su, Xiaosu, Sun, Zihan, Jia, Peilei, Gao, Jun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Cross-Utterance Conditioned VAE for Speech Generation
por: Li, Yang, et al.
Publicado: (2023)
por: Li, Yang, et al.
Publicado: (2023)
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
por: Liu, Zhanxun, et al.
Publicado: (2025)
por: Liu, Zhanxun, et al.
Publicado: (2025)
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
por: Zhao, Yiyang, et al.
Publicado: (2025)
por: Zhao, Yiyang, et al.
Publicado: (2025)
Hello-Chat: Towards Realistic Social Audio Interactions
por: Hou, Yueran, et al.
Publicado: (2026)
por: Hou, Yueran, et al.
Publicado: (2026)
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
por: Zhang, Xueyao, et al.
Publicado: (2025)
por: Zhang, Xueyao, et al.
Publicado: (2025)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
por: Zheng, Zhisheng, et al.
Publicado: (2025)
por: Zheng, Zhisheng, et al.
Publicado: (2025)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
por: Cheng, Sitong, et al.
Publicado: (2025)
por: Cheng, Sitong, et al.
Publicado: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
por: Anastassiou, Philip, et al.
Publicado: (2024)
por: Anastassiou, Philip, et al.
Publicado: (2024)
CapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation
por: Chu, Xuangeng, et al.
Publicado: (2026)
por: Chu, Xuangeng, et al.
Publicado: (2026)
TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis
por: Liang, Qifan, et al.
Publicado: (2026)
por: Liang, Qifan, et al.
Publicado: (2026)
Multi-Utterance Speech Separation and Association Trained on Short Segments
por: Wang, Yuzhu, et al.
Publicado: (2025)
por: Wang, Yuzhu, et al.
Publicado: (2025)
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
por: Flynn, Robert, et al.
Publicado: (2026)
por: Flynn, Robert, et al.
Publicado: (2026)
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
por: Wang, Yuzhu, et al.
Publicado: (2025)
por: Wang, Yuzhu, et al.
Publicado: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
por: Li, Hebeizi, et al.
Publicado: (2026)
por: Li, Hebeizi, et al.
Publicado: (2026)
On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation
por: Cheng, Changhao, et al.
Publicado: (2026)
por: Cheng, Changhao, et al.
Publicado: (2026)
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
por: Xue, Jun, et al.
Publicado: (2026)
por: Xue, Jun, et al.
Publicado: (2026)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
por: Huo, Mingyue, et al.
Publicado: (2025)
por: Huo, Mingyue, et al.
Publicado: (2025)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
por: Zhang, Enshi, et al.
Publicado: (2024)
por: Zhang, Enshi, et al.
Publicado: (2024)
VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis
por: Ma, Chengyuan, et al.
Publicado: (2026)
por: Ma, Chengyuan, et al.
Publicado: (2026)
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
por: Zhou, Yixuan, et al.
Publicado: (2025)
por: Zhou, Yixuan, et al.
Publicado: (2025)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
por: Zink, Oswald, et al.
Publicado: (2024)
por: Zink, Oswald, et al.
Publicado: (2024)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
por: Huang, Kuan-Po, et al.
Publicado: (2023)
por: Huang, Kuan-Po, et al.
Publicado: (2023)
Adapting Speech Language Model to Singing Voice Synthesis
por: Zhao, Yiwen, et al.
Publicado: (2025)
por: Zhao, Yiwen, et al.
Publicado: (2025)
Perpetual Dialogues: A Computational Analysis of Voice-Guitar Interaction in Carlos Paredes's Discography
por: Bernardes, Gilberto, et al.
Publicado: (2026)
por: Bernardes, Gilberto, et al.
Publicado: (2026)
VoiceBridge: General Speech Restoration with One-step Latent Bridge Models
por: Zhang, Chi, et al.
Publicado: (2025)
por: Zhang, Chi, et al.
Publicado: (2025)
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
por: Yang, Runyan, et al.
Publicado: (2024)
por: Yang, Runyan, et al.
Publicado: (2024)
Cross-Talk Speech Reduction, by Separation, for Separation
por: Wang, Zhong-Qiu, et al.
Publicado: (2026)
por: Wang, Zhong-Qiu, et al.
Publicado: (2026)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
por: Xie, Hanke, et al.
Publicado: (2025)
por: Xie, Hanke, et al.
Publicado: (2025)
DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions
por: Niu, Shu-Tong, et al.
Publicado: (2024)
por: Niu, Shu-Tong, et al.
Publicado: (2024)
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
por: Li, Fengjin, et al.
Publicado: (2025)
por: Li, Fengjin, et al.
Publicado: (2025)
VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
por: Zhou, Yixuan, et al.
Publicado: (2024)
por: Zhou, Yixuan, et al.
Publicado: (2024)
Speech to Speech Synthesis for Voice Impersonation
por: Johnson, Bjorn, et al.
Publicado: (2026)
por: Johnson, Bjorn, et al.
Publicado: (2026)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
por: Wang, Helin, et al.
Publicado: (2025)
por: Wang, Helin, et al.
Publicado: (2025)
DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
por: Shi, Weiyan, et al.
Publicado: (2025)
por: Shi, Weiyan, et al.
Publicado: (2025)
EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
por: Xu, Haiying, et al.
Publicado: (2025)
por: Xu, Haiying, et al.
Publicado: (2025)
Improving Short Utterance Anti-Spoofing with AASIST2
por: Zhang, Yuxiang, et al.
Publicado: (2023)
por: Zhang, Yuxiang, et al.
Publicado: (2023)
Geneses: Unified Generative Speech Enhancement and Separation
por: Asai, Kohei, et al.
Publicado: (2026)
por: Asai, Kohei, et al.
Publicado: (2026)
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
por: Kang, Fang, et al.
Publicado: (2025)
por: Kang, Fang, et al.
Publicado: (2025)
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
por: Guragain, Anmol, et al.
Publicado: (2024)
por: Guragain, Anmol, et al.
Publicado: (2024)
Ejemplares similares
-
Cross-Utterance Conditioned VAE for Speech Generation
por: Li, Yang, et al.
Publicado: (2023) -
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
por: Liu, Zhanxun, et al.
Publicado: (2025) -
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
por: Zhao, Yiyang, et al.
Publicado: (2025) -
Hello-Chat: Towards Realistic Social Audio Interactions
por: Hou, Yueran, et al.
Publicado: (2026) -
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
por: Zhang, Xueyao, et al.
Publicado: (2025)