Gespeichert in:
| Hauptverfasser: | Lee, Jaejun, Oh, Yoori, Lee, Kyogu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.01879 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency
von: Lee, Jaejun, et al.
Veröffentlicht: (2026)
von: Lee, Jaejun, et al.
Veröffentlicht: (2026)
Hear Your Face: Face-based voice conversion with F0 estimation
von: Lee, Jaejun, et al.
Veröffentlicht: (2024)
von: Lee, Jaejun, et al.
Veröffentlicht: (2024)
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
von: Lee, Jaejun, et al.
Veröffentlicht: (2025)
von: Lee, Jaejun, et al.
Veröffentlicht: (2025)
EMG-to-Speech with Fewer Channels
von: Hwang, Injune, et al.
Veröffentlicht: (2026)
von: Hwang, Injune, et al.
Veröffentlicht: (2026)
Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data Manipulation
von: Oh, Yoori, et al.
Veröffentlicht: (2024)
von: Oh, Yoori, et al.
Veröffentlicht: (2024)
Sounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection
von: Joo, Seohyun, et al.
Veröffentlicht: (2026)
von: Joo, Seohyun, et al.
Veröffentlicht: (2026)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
von: Han, Seungu, et al.
Veröffentlicht: (2026)
von: Han, Seungu, et al.
Veröffentlicht: (2026)
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling
von: Hwang, Injune, et al.
Veröffentlicht: (2024)
von: Hwang, Injune, et al.
Veröffentlicht: (2024)
Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement
von: Han, Seungu, et al.
Veröffentlicht: (2025)
von: Han, Seungu, et al.
Veröffentlicht: (2025)
String Sound Synthesizer on GPU-accelerated Finite Difference Scheme
von: Lee, Jin Woo, et al.
Veröffentlicht: (2023)
von: Lee, Jin Woo, et al.
Veröffentlicht: (2023)
Do Captioning Metrics Reflect Music Semantic Alignment?
von: Lee, Jinwoo, et al.
Veröffentlicht: (2024)
von: Lee, Jinwoo, et al.
Veröffentlicht: (2024)
MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
von: Lee, Jin Woo, et al.
Veröffentlicht: (2024)
von: Lee, Jin Woo, et al.
Veröffentlicht: (2024)
Music De-limiter Networks via Sample-wise Gain Inversion
von: Jeon, Chang-Bin, et al.
Veröffentlicht: (2023)
von: Jeon, Chang-Bin, et al.
Veröffentlicht: (2023)
Wavespace: A Highly Explorable Wavetable Generator
von: Lee, Hazounne, et al.
Veröffentlicht: (2024)
von: Lee, Hazounne, et al.
Veröffentlicht: (2024)
Music Auto-Tagging with Robust Music Representation Learned via Domain Adversarial Training
von: Joung, Haesun, et al.
Veröffentlicht: (2024)
von: Joung, Haesun, et al.
Veröffentlicht: (2024)
When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds
von: Kang, Minsu, et al.
Veröffentlicht: (2025)
von: Kang, Minsu, et al.
Veröffentlicht: (2025)
DDD: A Perceptually Superior Low-Response-Time DNN-based Declipper
von: Yi, Jayeon, et al.
Veröffentlicht: (2024)
von: Yi, Jayeon, et al.
Veröffentlicht: (2024)
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument
von: Kim, Kyungsu, et al.
Veröffentlicht: (2025)
von: Kim, Kyungsu, et al.
Veröffentlicht: (2025)
IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
von: Sankar, Ashwin, et al.
Veröffentlicht: (2024)
von: Sankar, Ashwin, et al.
Veröffentlicht: (2024)
Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis
von: Jeon, Yejin, et al.
Veröffentlicht: (2025)
von: Jeon, Yejin, et al.
Veröffentlicht: (2025)
VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
Inverse Nonlinearity Compensation of Hyperelastic Deformation in Dielectric Elastomer for Acoustic Actuation
von: Lee, Jin Woo, et al.
Veröffentlicht: (2024)
von: Lee, Jin Woo, et al.
Veröffentlicht: (2024)
Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
von: Kim, Eungbeom, et al.
Veröffentlicht: (2024)
von: Kim, Eungbeom, et al.
Veröffentlicht: (2024)
Speaking Clearly: A Simplified Whisper-Based Codec for Low-Bitrate Speech Coding
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2023)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2023)
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
von: Narain, Jaya, et al.
Veröffentlicht: (2025)
von: Narain, Jaya, et al.
Veröffentlicht: (2025)
A long-form single-speaker real-time MRI speech dataset and benchmark
von: Foley, Sean, et al.
Veröffentlicht: (2025)
von: Foley, Sean, et al.
Veröffentlicht: (2025)
Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
von: Lee, Myungjin, et al.
Veröffentlicht: (2026)
von: Lee, Myungjin, et al.
Veröffentlicht: (2026)
DOSE : Drum One-Shot Extraction from Music Mixture
von: Hwang, Suntae, et al.
Veröffentlicht: (2025)
von: Hwang, Suntae, et al.
Veröffentlicht: (2025)
VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
Searching For Music Mixing Graphs: A Pruning Approach
von: Lee, Sungho, et al.
Veröffentlicht: (2024)
von: Lee, Sungho, et al.
Veröffentlicht: (2024)
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
von: Pistrosch, Simon, et al.
Veröffentlicht: (2026)
von: Pistrosch, Simon, et al.
Veröffentlicht: (2026)
When Vision Speaks for Sound
von: Wen, Xiaofei, et al.
Veröffentlicht: (2026)
von: Wen, Xiaofei, et al.
Veröffentlicht: (2026)
Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
Practical and Reproducible Symbolic Music Generation by Large Language Models with Structural Embeddings
von: Rhyu, Seungyeon, et al.
Veröffentlicht: (2024)
von: Rhyu, Seungyeon, et al.
Veröffentlicht: (2024)
Hierarchical speaker representation for target speaker extraction
von: He, Shulin, et al.
Veröffentlicht: (2022)
von: He, Shulin, et al.
Veröffentlicht: (2022)
Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency
von: Lee, Jaejun, et al.
Veröffentlicht: (2026) -
Hear Your Face: Face-based voice conversion with F0 estimation
von: Lee, Jaejun, et al.
Veröffentlicht: (2024) -
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
von: Lee, Jaejun, et al.
Veröffentlicht: (2025) -
EMG-to-Speech with Fewer Channels
von: Hwang, Injune, et al.
Veröffentlicht: (2026) -
Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data Manipulation
von: Oh, Yoori, et al.
Veröffentlicht: (2024)