JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yu, Fan, Wang, Tao, Wu, You, Zhu, Lin, Deng, Wei, Han, Weisheng, Wang, Wenchao, Hu, Lin, Liang, Xiangyu, He, Xiaodong, Huang, Yankun, Gu, Yu, Liu, Yuan, Wang, Yuxuan, Xiao, Zhangyu, Wang, Ziteng, Dong, Boya, Dang, Feng, Chen, Jinming, Li, Jingdong, Wang, Jun, Jin, Yechen, Zhang, Yuan, Sheng, Zhengyan, Wang, Xin |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Generating Novel and Realistic Speakers for Voice Conversion
par: Chen, Meiying Melissa, et autres
Publié: (2025)
par: Chen, Meiying Melissa, et autres
Publié: (2025)
JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
par: Zhou, Fangru, et autres
Publié: (2025)
par: Zhou, Fangru, et autres
Publié: (2025)
Learning Equilibrium Fluctuation Expansions from Overdamped Langevin Dynamics
par: Wang, Lin, et autres
Publié: (2026)
par: Wang, Lin, et autres
Publié: (2026)
Probabilistic Approaches to The Energy Equality in Forced Surface Quasi-Geostrophic Equations
par: Wang, Lin, et autres
Publié: (2024)
par: Wang, Lin, et autres
Publié: (2024)
ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations
par: Wang, Kexue, et autres
Publié: (2026)
par: Wang, Kexue, et autres
Publié: (2026)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2023)
par: Wang, Zhichao, et autres
Publié: (2023)
Residual Speaker Representation for One-Shot Voice Conversion
par: Xu, Le, et autres
Publié: (2023)
par: Xu, Le, et autres
Publié: (2023)
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
par: Ren, Pengyu, et autres
Publié: (2025)
par: Ren, Pengyu, et autres
Publié: (2025)
Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning
par: He, Haorui, et autres
Publié: (2024)
par: He, Haorui, et autres
Publié: (2024)
Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
par: Cai, Runyuan, et autres
Publié: (2026)
par: Cai, Runyuan, et autres
Publié: (2026)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2024)
par: Wang, Zhichao, et autres
Publié: (2024)
Virasoro constraints for K3 surfaces and monodromy operators
par: Wang, Weisheng
Publié: (2024)
par: Wang, Weisheng
Publié: (2024)
Talking Spell: A Wearable System Enabling Real-Time Anthropomorphic Voice Interaction with Everyday Objects
par: Wang, Xuetong, et autres
Publié: (2025)
par: Wang, Xuetong, et autres
Publié: (2025)
Humanlike AI for Corporate Social Responsibility Communication: How Perceived Anthropomorphism Shapes Stakeholder Acceptance of Chatbots
par: Yangzhi (Nicole) Jiang, et autres
Publié: (2026)
par: Yangzhi (Nicole) Jiang, et autres
Publié: (2026)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2024)
par: Wang, Zhichao, et autres
Publié: (2024)
From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
par: Xu, Ke, et autres
Publié: (2026)
par: Xu, Ke, et autres
Publié: (2026)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
par: Hou, Yixuan, et autres
Publié: (2025)
par: Hou, Yixuan, et autres
Publié: (2025)
Human-Inspired Soft Anthropomorphic Hand System for Neuromorphic Object and Pose Recognition Using Multimodal Signals
par: Wang, Fengyi, et autres
Publié: (2025)
par: Wang, Fengyi, et autres
Publié: (2025)
Object Classification Utilizing Neuromorphic Proprioceptive Signals in Active Exploration: Validated on a Soft Anthropomorphic Hand
par: Wang, Fengyi, et autres
Publié: (2025)
par: Wang, Fengyi, et autres
Publié: (2025)
One-variable equations over the lamplighter group
par: Ushakov, Alexander, et autres
Publié: (2026)
par: Ushakov, Alexander, et autres
Publié: (2026)
One variable equations over the lamplighter group
par: Ushakov, Alexander, et autres
Publié: (2025)
par: Ushakov, Alexander, et autres
Publié: (2025)
Online Prediction of Operating Temperature for Permanent Magnet Motor Used in EVs Based on Parameter Identification
par: Wang Yankun, et autres
Publié: (2025)
par: Wang Yankun, et autres
Publié: (2025)
SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
par: Yin, Chun, et autres
Publié: (2024)
par: Yin, Chun, et autres
Publié: (2024)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
par: Yuan, Ying, et autres
Publié: (2026)
par: Yuan, Ying, et autres
Publié: (2026)
DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage
par: Wang, Kyra, et autres
Publié: (2024)
par: Wang, Kyra, et autres
Publié: (2024)
Underwater Acoustic Target Recognition based on Smoothness-inducing Regularization and Spectrogram-based Data Augmentation
par: Xu, Ji, et autres
Publié: (2023)
par: Xu, Ji, et autres
Publié: (2023)
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
par: You, Zhenghai, et autres
Publié: (2025)
par: You, Zhenghai, et autres
Publié: (2025)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
par: Wang, Haoyu, et autres
Publié: (2024)
par: Wang, Haoyu, et autres
Publié: (2024)
Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
par: Qi, Tianhua, et autres
Publié: (2024)
par: Qi, Tianhua, et autres
Publié: (2024)
JoyAgent-JDGenie: Technical Report on the GAIA
par: Liu, Jiarun, et autres
Publié: (2025)
par: Liu, Jiarun, et autres
Publié: (2025)
DreamVoice: Text-Guided Voice Conversion
par: Hai, Jiarui, et autres
Publié: (2024)
par: Hai, Jiarui, et autres
Publié: (2024)
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
par: Li, Junjie, et autres
Publié: (2023)
par: Li, Junjie, et autres
Publié: (2023)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
par: Zhao, Junchuan, et autres
Publié: (2025)
par: Zhao, Junchuan, et autres
Publié: (2025)
USM-VC: Mitigating Timbre Leakage with Universal Semantic Mapping Residual Block for Voice Conversion
par: Li, Na, et autres
Publié: (2025)
par: Li, Na, et autres
Publié: (2025)
Voice Conversion Augmentation for Speaker Recognition on Defective Datasets
par: Tao, Ruijie, et autres
Publié: (2024)
par: Tao, Ruijie, et autres
Publié: (2024)
What Does the Speaker Embedding Encode?
par: Wang, Shuai, et autres
Publié: (2025)
par: Wang, Shuai, et autres
Publié: (2025)
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
par: Wang, Ziteng, et autres
Publié: (2025)
par: Wang, Ziteng, et autres
Publié: (2025)
Strong limit theorems / Lin Zhengyan and Lu Chuanrong
par: Zhengyan, Lin
par: Zhengyan, Lin
Any-to-any Speaker Attribute Perturbation for Asynchronous Voice Anonymization
par: Chen, Liping, et autres
Publié: (2025)
par: Chen, Liping, et autres
Publié: (2025)
Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
par: Wang, Rui, et autres
Publié: (2024)
par: Wang, Rui, et autres
Publié: (2024)
Documents similaires
-
Generating Novel and Realistic Speakers for Voice Conversion
par: Chen, Meiying Melissa, et autres
Publié: (2025) -
JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
par: Zhou, Fangru, et autres
Publié: (2025) -
Learning Equilibrium Fluctuation Expansions from Overdamped Langevin Dynamics
par: Wang, Lin, et autres
Publié: (2026) -
Probabilistic Approaches to The Energy Equality in Forced Surface Quasi-Geostrophic Equations
par: Wang, Lin, et autres
Publié: (2024) -
ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations
par: Wang, Kexue, et autres
Publié: (2026)