DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Oh, Hyung-Seok, Lee, Sang-Hoon, Cho, Deok-Hyeon, Lee, Seong-Whan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2026)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2026)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
Toward Complex-Valued Neural Networks for Waveform Generation
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2026)
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2026)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
von: Cha, Jun-Hyeok, et al.
Veröffentlicht: (2025)
von: Cha, Jun-Hyeok, et al.
Veröffentlicht: (2025)
VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion
von: Choi, Joon-Seung, et al.
Veröffentlicht: (2025)
von: Choi, Joon-Seung, et al.
Veröffentlicht: (2025)
TranSentence: Speech-to-speech Translation via Language-agnostic Sentence-level Speech Encoding without Language-parallel Data
von: Kim, Seung-Bin, et al.
Veröffentlicht: (2024)
von: Kim, Seung-Bin, et al.
Veröffentlicht: (2024)
PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
von: Qi, Tianhua, et al.
Veröffentlicht: (2025)
von: Qi, Tianhua, et al.
Veröffentlicht: (2025)
Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
von: Yoon, Jinsung, et al.
Veröffentlicht: (2025)
von: Yoon, Jinsung, et al.
Veröffentlicht: (2025)
SUGAR: Leveraging Contextual Confidence for Smarter Retrieval
von: Zubkova, Hanna, et al.
Veröffentlicht: (2025)
von: Zubkova, Hanna, et al.
Veröffentlicht: (2025)
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
Calibration-Free EEG-based Driver Drowsiness Detection with Online Test-Time Adaptation
von: Jang, Geun-Deok, et al.
Veröffentlicht: (2025)
von: Jang, Geun-Deok, et al.
Veröffentlicht: (2025)
KiC: Keyword-inspired Cascade for Cost-Efficient Text Generation with LLMs
von: Kim, Woo-Chan, et al.
Veröffentlicht: (2025)
von: Kim, Woo-Chan, et al.
Veröffentlicht: (2025)
Reconstructing Unseen Sentences from Speech-related Biosignals for Open-vocabulary Neural Communication
von: Kim, Deok-Seon, et al.
Veröffentlicht: (2025)
von: Kim, Deok-Seon, et al.
Veröffentlicht: (2025)
Personalized targeted memory reactivation enhances consolidation of challenging memories via slow wave and spindle dynamics
von: Shin, Gi-Hwan, et al.
Veröffentlicht: (2025)
von: Shin, Gi-Hwan, et al.
Veröffentlicht: (2025)
DQE-CIR: Distinctive Query Embeddings through Learnable Attribute Weights and Target Relative Negative Sampling in Composed Image Retrieval
von: Park, Geon, et al.
Veröffentlicht: (2026)
von: Park, Geon, et al.
Veröffentlicht: (2026)
Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition
von: Jung, Gunho, et al.
Veröffentlicht: (2025)
von: Jung, Gunho, et al.
Veröffentlicht: (2025)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
von: Kim, Ji-Hyeon, et al.
Veröffentlicht: (2026)
von: Kim, Ji-Hyeon, et al.
Veröffentlicht: (2026)
Local Representative Token Guided Merging for Text-to-Image Generation
von: Lee, Min-Jeong, et al.
Veröffentlicht: (2025)
von: Lee, Min-Jeong, et al.
Veröffentlicht: (2025)
Sleep‐enhancing activity of fermented pea protein hydrolysate with enhanced GABA content by Lactobacillus brevisSYLB 0016 fermentation
von: Hyeon Deok Kim, et al.
Veröffentlicht: (2025)
von: Hyeon Deok Kim, et al.
Veröffentlicht: (2025)
Visual Motion Imagery Classification with Deep Neural Network based on Functional Connectivity
von: Kwon, Byoung-Hee, et al.
Veröffentlicht: (2021)
von: Kwon, Byoung-Hee, et al.
Veröffentlicht: (2021)
Explaining generative diffusion models via visual analysis for interpretable decision-making process
von: Park, Ji-Hoon, et al.
Veröffentlicht: (2024)
von: Park, Ji-Hoon, et al.
Veröffentlicht: (2024)
Ternary Logic Transistors Using Multi‐Stacked 2D Electron Gas Channels in Ultrathin Oxide Heterostructures
von: Ji Hyeon Choi, et al.
Veröffentlicht: (2024)
von: Ji Hyeon Choi, et al.
Veröffentlicht: (2024)
Ternary Logic Transistors Using Multi‐Stacked 2D Electron Gas Channels in Ultrathin Oxide Heterostructures (Adv. Sci. 6/2025)
von: Ji Hyeon Choi, et al.
Veröffentlicht: (2025)
von: Ji Hyeon Choi, et al.
Veröffentlicht: (2025)
PVDF ‐Pyrolyzed Fluorine‐Doped TiO 2 for Synergistic Adsorption‐Enhanced PFOA Photocatalysis
von: Deok Hoon Kim, et al.
Veröffentlicht: (2026)
von: Deok Hoon Kim, et al.
Veröffentlicht: (2026)
Influence of carcass mass on decomposition rate: A medico‐legal entomology perspective
von: Hyeon‐Seok Oh, et al.
Veröffentlicht: (2024)
von: Hyeon‐Seok Oh, et al.
Veröffentlicht: (2024)
ID-EA: Identity-driven Text Enhancement and Adaptation with Textual Inversion for Personalized Text-to-Image Generation
von: Jin, Hyun-Jun, et al.
Veröffentlicht: (2025)
von: Jin, Hyun-Jun, et al.
Veröffentlicht: (2025)
Seasonal variation in risk and return trade‐off
von: Deok‐Hyeon Lee, et al.
Veröffentlicht: (2024)
von: Deok‐Hyeon Lee, et al.
Veröffentlicht: (2024)
EEG-based Multimodal Representation Learning for Emotion Recognition
von: Yin, Kang, et al.
Veröffentlicht: (2024)
von: Yin, Kang, et al.
Veröffentlicht: (2024)
Edge Conditional Node Update Graph Neural Network for Multi-variate Time Series Anomaly Detection
von: Jo, Hayoung, et al.
Veröffentlicht: (2024)
von: Jo, Hayoung, et al.
Veröffentlicht: (2024)
TIFu: Tri-directional Implicit Function for High-Fidelity 3D Character Reconstruction
von: Lim, Byoungsung, et al.
Veröffentlicht: (2024)
von: Lim, Byoungsung, et al.
Veröffentlicht: (2024)
Influence of Video Dynamics on EEG-based Single-Trial Video Target Surveillance System
von: Kwak, Heon-Gyu, et al.
Veröffentlicht: (2023)
von: Kwak, Heon-Gyu, et al.
Veröffentlicht: (2023)
Quantitative Coronary Angiography Guidance for Drug‐Eluting Stent Implantation: A Narrative Review
von: Cheol Whan Lee, et al.
Veröffentlicht: (2024)
von: Cheol Whan Lee, et al.
Veröffentlicht: (2024)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
von: Oh, Ju-Young, et al.
Veröffentlicht: (2025)
von: Oh, Ju-Young, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024) -
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024) -
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025) -
Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2026) -
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)