LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Jaejun, Oh, Yoori, Lee, Kyogu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
di: Lee, Jaejun, et al.
Pubblicazione: (2026)
di: Lee, Jaejun, et al.
Pubblicazione: (2026)
Hear Your Face: Face-based voice conversion with F0 estimation
di: Lee, Jaejun, et al.
Pubblicazione: (2024)
di: Lee, Jaejun, et al.
Pubblicazione: (2024)
EMG-to-Speech with Fewer Channels
di: Hwang, Injune, et al.
Pubblicazione: (2026)
di: Hwang, Injune, et al.
Pubblicazione: (2026)
Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data Manipulation
di: Oh, Yoori, et al.
Pubblicazione: (2024)
di: Oh, Yoori, et al.
Pubblicazione: (2024)
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
di: Lee, Jaejun, et al.
Pubblicazione: (2025)
di: Lee, Jaejun, et al.
Pubblicazione: (2025)
LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
di: Richter, Julius, et al.
Pubblicazione: (2025)
di: Richter, Julius, et al.
Pubblicazione: (2025)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023)
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023)
LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading
di: Yemini, Yochai, et al.
Pubblicazione: (2023)
di: Yemini, Yochai, et al.
Pubblicazione: (2023)
Towards Accurate Lip-to-Speech Synthesis in-the-Wild
di: Hegde, Sindhu, et al.
Pubblicazione: (2024)
di: Hegde, Sindhu, et al.
Pubblicazione: (2024)
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
di: Park, Young-Hu, et al.
Pubblicazione: (2025)
di: Park, Young-Hu, et al.
Pubblicazione: (2025)
A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync Synthesis
di: Amir, Javeria, et al.
Pubblicazione: (2025)
di: Amir, Javeria, et al.
Pubblicazione: (2025)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
di: Han, Seungu, et al.
Pubblicazione: (2026)
di: Han, Seungu, et al.
Pubblicazione: (2026)
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
di: Chae, Yunkee, et al.
Pubblicazione: (2025)
di: Chae, Yunkee, et al.
Pubblicazione: (2025)
Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement
di: Han, Seungu, et al.
Pubblicazione: (2025)
di: Han, Seungu, et al.
Pubblicazione: (2025)
Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling
di: Hwang, Injune, et al.
Pubblicazione: (2024)
di: Hwang, Injune, et al.
Pubblicazione: (2024)
LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning
di: Yang, Kang, et al.
Pubblicazione: (2025)
di: Yang, Kang, et al.
Pubblicazione: (2025)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
di: Goncalves, Lucas, et al.
Pubblicazione: (2024)
di: Goncalves, Lucas, et al.
Pubblicazione: (2024)
Do Captioning Metrics Reflect Music Semantic Alignment?
di: Lee, Jinwoo, et al.
Pubblicazione: (2024)
di: Lee, Jinwoo, et al.
Pubblicazione: (2024)
Sounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection
di: Joo, Seohyun, et al.
Pubblicazione: (2026)
di: Joo, Seohyun, et al.
Pubblicazione: (2026)
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
di: Li, Kai, et al.
Pubblicazione: (2025)
di: Li, Kai, et al.
Pubblicazione: (2025)
NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing
di: Liang, Yifan, et al.
Pubblicazione: (2025)
di: Liang, Yifan, et al.
Pubblicazione: (2025)
Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder
di: Dai, Yusheng, et al.
Pubblicazione: (2023)
di: Dai, Yusheng, et al.
Pubblicazione: (2023)
MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction
di: Chae, Yunkee, et al.
Pubblicazione: (2025)
di: Chae, Yunkee, et al.
Pubblicazione: (2025)
Music De-limiter Networks via Sample-wise Gain Inversion
di: Jeon, Chang-Bin, et al.
Pubblicazione: (2023)
di: Jeon, Chang-Bin, et al.
Pubblicazione: (2023)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
di: Jiang, Yuepeng, et al.
Pubblicazione: (2024)
di: Jiang, Yuepeng, et al.
Pubblicazione: (2024)
Wavespace: A Highly Explorable Wavetable Generator
di: Lee, Hazounne, et al.
Pubblicazione: (2024)
di: Lee, Hazounne, et al.
Pubblicazione: (2024)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
di: Gu, Ke, et al.
Pubblicazione: (2025)
di: Gu, Ke, et al.
Pubblicazione: (2025)
Music Auto-Tagging with Robust Music Representation Learned via Domain Adversarial Training
di: Joung, Haesun, et al.
Pubblicazione: (2024)
di: Joung, Haesun, et al.
Pubblicazione: (2024)
Sign-to-Speech Prosody Transfer via Sign Reconstruction-based GAN
di: Manabe, Toranosuke, et al.
Pubblicazione: (2026)
di: Manabe, Toranosuke, et al.
Pubblicazione: (2026)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
di: He, Xiangheng, et al.
Pubblicazione: (2024)
di: He, Xiangheng, et al.
Pubblicazione: (2024)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
di: Du, Chenpeng, et al.
Pubblicazione: (2021)
di: Du, Chenpeng, et al.
Pubblicazione: (2021)
Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2026)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2026)
EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning
di: Wang, Dingdong, et al.
Pubblicazione: (2026)
di: Wang, Dingdong, et al.
Pubblicazione: (2026)
Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
di: Lee, Jin Woo, et al.
Pubblicazione: (2024)
di: Lee, Jin Woo, et al.
Pubblicazione: (2024)
TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument
di: Kim, Kyungsu, et al.
Pubblicazione: (2025)
di: Kim, Kyungsu, et al.
Pubblicazione: (2025)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
di: Koriyama, Tomoki
Pubblicazione: (2025)
di: Koriyama, Tomoki
Pubblicazione: (2025)
LipsAM: Lipschitz-Continuous Amplitude Modifier for Audio Signal Processing and its Application to Plug-and-Play Dereverberation
di: Matsumoto, Kazuki, et al.
Pubblicazione: (2026)
di: Matsumoto, Kazuki, et al.
Pubblicazione: (2026)
DDD: A Perceptually Superior Low-Response-Time DNN-based Declipper
di: Yi, Jayeon, et al.
Pubblicazione: (2024)
di: Yi, Jayeon, et al.
Pubblicazione: (2024)
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
String Sound Synthesizer on GPU-accelerated Finite Difference Scheme
di: Lee, Jin Woo, et al.
Pubblicazione: (2023)
di: Lee, Jin Woo, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
di: Lee, Jaejun, et al.
Pubblicazione: (2026) -
Hear Your Face: Face-based voice conversion with F0 estimation
di: Lee, Jaejun, et al.
Pubblicazione: (2024) -
EMG-to-Speech with Fewer Channels
di: Hwang, Injune, et al.
Pubblicazione: (2026) -
Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data Manipulation
di: Oh, Yoori, et al.
Pubblicazione: (2024) -
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
di: Lee, Jaejun, et al.
Pubblicazione: (2025)