Text-Driven Voice Conversion via Latent State-Space Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Wen, Martinez, Sofia, Shah, Priyanka |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DGFM: Full Body Dance Generation Driven by Music Foundation Models
by: Liu, Xinran, et al.
Published: (2025)
by: Liu, Xinran, et al.
Published: (2025)
Content and Style Aware Audio-Driven Facial Animation
by: Liu, Qingju, et al.
Published: (2024)
by: Liu, Qingju, et al.
Published: (2024)
EnchantDance: Unveiling the Potential of Music-Driven Dance Movement
by: Han, Bo, et al.
Published: (2023)
by: Han, Bo, et al.
Published: (2023)
NAT: Neural Acoustic Transfer for Interactive Scenes in Real Time
by: Jin, Xutong, et al.
Published: (2025)
by: Jin, Xutong, et al.
Published: (2025)
ELGAR: Expressive Cello Performance Motion Generation for Audio Rendition
by: Qiu, Zhiping, et al.
Published: (2025)
by: Qiu, Zhiping, et al.
Published: (2025)
Listen and Move: Improving GANs Coherency in Agnostic Sound-to-Video Generation
by: Redondo, Rafael
Published: (2024)
by: Redondo, Rafael
Published: (2024)
SyncViolinist: Music-Oriented Violin Motion Generation Based on Bowing and Fingering
by: Nishizawa, Hiroki, et al.
Published: (2024)
by: Nishizawa, Hiroki, et al.
Published: (2024)
Beat-It: Beat-Synchronized Multi-Condition 3D Dance Generation
by: Huang, Zikai, et al.
Published: (2024)
by: Huang, Zikai, et al.
Published: (2024)
Combining Genre Classification and Harmonic-Percussive Features with Diffusion Models for Music-Video Generation
by: Pina, Leonardo, et al.
Published: (2024)
by: Pina, Leonardo, et al.
Published: (2024)
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
by: Kameoka, Hirokazu, et al.
Published: (2025)
by: Kameoka, Hirokazu, et al.
Published: (2025)
MusicScore: A Dataset for Music Score Modeling and Generation
by: Lin, Yuheng, et al.
Published: (2024)
by: Lin, Yuheng, et al.
Published: (2024)
Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
by: Zhou, Xukun, et al.
Published: (2024)
by: Zhou, Xukun, et al.
Published: (2024)
MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation
by: Yang, Kaixing, et al.
Published: (2025)
by: Yang, Kaixing, et al.
Published: (2025)
DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning
by: Bhattacharya, Aneesh, et al.
Published: (2023)
by: Bhattacharya, Aneesh, et al.
Published: (2023)
LCM-SVC: Latent Diffusion Model Based Singing Voice Conversion with Inference Acceleration via Latent Consistency Distillation
by: Chen, Shihao, et al.
Published: (2024)
by: Chen, Shihao, et al.
Published: (2024)
Gaunt coefficients for complex and real spherical harmonics with applications to spherical array processing and Ambisonics
by: Politis, Archontis
Published: (2024)
by: Politis, Archontis
Published: (2024)
DuetGen: Music Driven Two-Person Dance Generation via Hierarchical Masked Modeling
by: Ghosh, Anindita, et al.
Published: (2025)
by: Ghosh, Anindita, et al.
Published: (2025)
TCDiff++: An End-to-end Trajectory-Controllable Diffusion Model for Harmonious Music-Driven Group Choreography
by: Dai, Yuqin, et al.
Published: (2025)
by: Dai, Yuqin, et al.
Published: (2025)
ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training
by: Zhu, Xinfa, et al.
Published: (2025)
by: Zhu, Xinfa, et al.
Published: (2025)
LAV: Audio-Driven Dynamic Visual Generation with Neural Compression and StyleGAN2
by: Jung, Jongmin, et al.
Published: (2025)
by: Jung, Jongmin, et al.
Published: (2025)
Text-driven Talking Face Synthesis by Reprogramming Audio-driven Models
by: Choi, Jeongsoo, et al.
Published: (2023)
by: Choi, Jeongsoo, et al.
Published: (2023)
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
by: Xie, Yifan, et al.
Published: (2024)
by: Xie, Yifan, et al.
Published: (2024)
LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
by: Chen, Shihao, et al.
Published: (2024)
by: Chen, Shihao, et al.
Published: (2024)
FürElise: Capturing and Physically Synthesizing Hand Motions of Piano Performance
by: Wang, Ruocheng, et al.
Published: (2024)
by: Wang, Ruocheng, et al.
Published: (2024)
Open Your Ears and Take a Look: A State-of-the-Art Report on the Integration of Sonification and Visualization
by: Enge, Kajetan, et al.
Published: (2024)
by: Enge, Kajetan, et al.
Published: (2024)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
by: Gudmalwar, Ashishkumar, et al.
Published: (2024)
by: Gudmalwar, Ashishkumar, et al.
Published: (2024)
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
by: Wang, Zhichao, et al.
Published: (2026)
by: Wang, Zhichao, et al.
Published: (2026)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
by: Li, Yuke, et al.
Published: (2024)
by: Li, Yuke, et al.
Published: (2024)
Residual Speaker Representation for One-Shot Voice Conversion
by: Xu, Le, et al.
Published: (2023)
by: Xu, Le, et al.
Published: (2023)
Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion
by: Sha, Binzhu, et al.
Published: (2023)
by: Sha, Binzhu, et al.
Published: (2023)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
by: Seki, Kentaro, et al.
Published: (2024)
by: Seki, Kentaro, et al.
Published: (2024)
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
by: Igarashi, Takuto, et al.
Published: (2024)
by: Igarashi, Takuto, et al.
Published: (2024)
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
by: Gudmalwar, Ashishkumar, et al.
Published: (2024)
by: Gudmalwar, Ashishkumar, et al.
Published: (2024)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
by: Du, Zongyang, et al.
Published: (2024)
by: Du, Zongyang, et al.
Published: (2024)
Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation
by: Huang, Wen, et al.
Published: (2025)
by: Huang, Wen, et al.
Published: (2025)
Audio is all in one: speech-driven gesture synthetics using WavLM pre-trained model
by: Zhang, Fan, et al.
Published: (2023)
by: Zhang, Fan, et al.
Published: (2023)
GCDance: Genre-Controlled Music-Driven 3D Full Body Dance Generation
by: Liu, Xinran, et al.
Published: (2025)
by: Liu, Xinran, et al.
Published: (2025)
Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models
by: Kim, Heeseung, et al.
Published: (2025)
by: Kim, Heeseung, et al.
Published: (2025)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
by: Wang, Zhichao, et al.
Published: (2024)
by: Wang, Zhichao, et al.
Published: (2024)
An Extensive Analysis of the Singing Voice Conversion Challenge 2025 Evaluation Results
by: Violeta, Lester Phillip, et al.
Published: (2025)
by: Violeta, Lester Phillip, et al.
Published: (2025)
Similar Items
-
DGFM: Full Body Dance Generation Driven by Music Foundation Models
by: Liu, Xinran, et al.
Published: (2025) -
Content and Style Aware Audio-Driven Facial Animation
by: Liu, Qingju, et al.
Published: (2024) -
EnchantDance: Unveiling the Potential of Music-Driven Dance Movement
by: Han, Bo, et al.
Published: (2023) -
NAT: Neural Acoustic Transfer for Interactive Scenes in Real Time
by: Jin, Xutong, et al.
Published: (2025) -
ELGAR: Expressive Cello Performance Motion Generation for Audio Rendition
by: Qiu, Zhiping, et al.
Published: (2025)