Saved in:
| Main Authors: | Lee, Jaejun, Oh, Yoori, Lee, Kyogu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.01908 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
by: Lee, Jaejun, et al.
Published: (2026)
by: Lee, Jaejun, et al.
Published: (2026)
Hear Your Face: Face-based voice conversion with F0 estimation
by: Lee, Jaejun, et al.
Published: (2024)
by: Lee, Jaejun, et al.
Published: (2024)
EMG-to-Speech with Fewer Channels
by: Hwang, Injune, et al.
Published: (2026)
by: Hwang, Injune, et al.
Published: (2026)
Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data Manipulation
by: Oh, Yoori, et al.
Published: (2024)
by: Oh, Yoori, et al.
Published: (2024)
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
by: Lee, Jaejun, et al.
Published: (2025)
by: Lee, Jaejun, et al.
Published: (2025)
LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
by: Richter, Julius, et al.
Published: (2025)
by: Richter, Julius, et al.
Published: (2025)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
by: Oh, Hyung-Seok, et al.
Published: (2023)
by: Oh, Hyung-Seok, et al.
Published: (2023)
LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading
by: Yemini, Yochai, et al.
Published: (2023)
by: Yemini, Yochai, et al.
Published: (2023)
Towards Accurate Lip-to-Speech Synthesis in-the-Wild
by: Hegde, Sindhu, et al.
Published: (2024)
by: Hegde, Sindhu, et al.
Published: (2024)
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
by: Park, Young-Hu, et al.
Published: (2025)
by: Park, Young-Hu, et al.
Published: (2025)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
by: Han, Seungu, et al.
Published: (2026)
by: Han, Seungu, et al.
Published: (2026)
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
by: Chae, Yunkee, et al.
Published: (2025)
by: Chae, Yunkee, et al.
Published: (2025)
A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync Synthesis
by: Amir, Javeria, et al.
Published: (2025)
by: Amir, Javeria, et al.
Published: (2025)
Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling
by: Hwang, Injune, et al.
Published: (2024)
by: Hwang, Injune, et al.
Published: (2024)
Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement
by: Han, Seungu, et al.
Published: (2025)
by: Han, Seungu, et al.
Published: (2025)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
by: Goncalves, Lucas, et al.
Published: (2024)
by: Goncalves, Lucas, et al.
Published: (2024)
LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
Do Captioning Metrics Reflect Music Semantic Alignment?
by: Lee, Jinwoo, et al.
Published: (2024)
by: Lee, Jinwoo, et al.
Published: (2024)
Sounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection
by: Joo, Seohyun, et al.
Published: (2026)
by: Joo, Seohyun, et al.
Published: (2026)
MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction
by: Chae, Yunkee, et al.
Published: (2025)
by: Chae, Yunkee, et al.
Published: (2025)
Music De-limiter Networks via Sample-wise Gain Inversion
by: Jeon, Chang-Bin, et al.
Published: (2023)
by: Jeon, Chang-Bin, et al.
Published: (2023)
Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder
by: Dai, Yusheng, et al.
Published: (2023)
by: Dai, Yusheng, et al.
Published: (2023)
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
by: Li, Kai, et al.
Published: (2025)
by: Li, Kai, et al.
Published: (2025)
NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing
by: Liang, Yifan, et al.
Published: (2025)
by: Liang, Yifan, et al.
Published: (2025)
Wavespace: A Highly Explorable Wavetable Generator
by: Lee, Hazounne, et al.
Published: (2024)
by: Lee, Hazounne, et al.
Published: (2024)
Music Auto-Tagging with Robust Music Representation Learned via Domain Adversarial Training
by: Joung, Haesun, et al.
Published: (2024)
by: Joung, Haesun, et al.
Published: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
by: Jiang, Yuepeng, et al.
Published: (2024)
by: Jiang, Yuepeng, et al.
Published: (2024)
Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
by: Lee, Jin Woo, et al.
Published: (2024)
by: Lee, Jin Woo, et al.
Published: (2024)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
by: Gu, Ke, et al.
Published: (2025)
by: Gu, Ke, et al.
Published: (2025)
DDD: A Perceptually Superior Low-Response-Time DNN-based Declipper
by: Yi, Jayeon, et al.
Published: (2024)
by: Yi, Jayeon, et al.
Published: (2024)
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
by: Kim, Jaeyeon, et al.
Published: (2024)
by: Kim, Jaeyeon, et al.
Published: (2024)
TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument
by: Kim, Kyungsu, et al.
Published: (2025)
by: Kim, Kyungsu, et al.
Published: (2025)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
by: He, Xiangheng, et al.
Published: (2024)
by: He, Xiangheng, et al.
Published: (2024)
String Sound Synthesizer on GPU-accelerated Finite Difference Scheme
by: Lee, Jin Woo, et al.
Published: (2023)
by: Lee, Jin Woo, et al.
Published: (2023)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
by: Du, Chenpeng, et al.
Published: (2021)
by: Du, Chenpeng, et al.
Published: (2021)
Sign-to-Speech Prosody Transfer via Sign Reconstruction-based GAN
by: Manabe, Toranosuke, et al.
Published: (2026)
by: Manabe, Toranosuke, et al.
Published: (2026)
Inverse Nonlinearity Compensation of Hyperelastic Deformation in Dielectric Elastomer for Acoustic Actuation
by: Lee, Jin Woo, et al.
Published: (2024)
by: Lee, Jin Woo, et al.
Published: (2024)
Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
by: Cho, Deok-Hyeon, et al.
Published: (2026)
by: Cho, Deok-Hyeon, et al.
Published: (2026)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
by: Karapiperis, Sotirios, et al.
Published: (2024)
by: Karapiperis, Sotirios, et al.
Published: (2024)
Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
by: Kim, Eungbeom, et al.
Published: (2024)
by: Kim, Eungbeom, et al.
Published: (2024)
Similar Items
-
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
by: Lee, Jaejun, et al.
Published: (2026) -
Hear Your Face: Face-based voice conversion with F0 estimation
by: Lee, Jaejun, et al.
Published: (2024) -
EMG-to-Speech with Fewer Channels
by: Hwang, Injune, et al.
Published: (2026) -
Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data Manipulation
by: Oh, Yoori, et al.
Published: (2024) -
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
by: Lee, Jaejun, et al.
Published: (2025)