RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Pham, Long-Khanh, Tran, Thanh V. T., Pham, Minh-Tan, Nguyen, Van |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
par: Pham, The Hieu, et autres
Publié: (2025)
par: Pham, The Hieu, et autres
Publié: (2025)
Emotional Vietnamese Speech-Based Depression Diagnosis Using Dynamic Attention Mechanism
par: D., Quang-Anh N., et autres
Publié: (2024)
par: D., Quang-Anh N., et autres
Publié: (2024)
Optimal Scalogram for Computational Complexity Reduction in Acoustic Recognition Using Deep Learning
par: Phan, Dang Thoai, et autres
Publié: (2025)
par: Phan, Dang Thoai, et autres
Publié: (2025)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
par: Le, Khanh, et autres
Publié: (2025)
par: Le, Khanh, et autres
Publié: (2025)
Shushing! Let's Imagine an Authentic Speech from the Silent Video
par: Ye, Jiaxin, et autres
Publié: (2025)
par: Ye, Jiaxin, et autres
Publié: (2025)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
par: Le, Khanh, et autres
Publié: (2025)
par: Le, Khanh, et autres
Publié: (2025)
FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
par: Zhang, Yiming, et autres
Publié: (2024)
par: Zhang, Yiming, et autres
Publié: (2024)
Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization
par: Ho, Luong, et autres
Publié: (2025)
par: Ho, Luong, et autres
Publié: (2025)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
par: Le-Duc, Khai, et autres
Publié: (2024)
par: Le-Duc, Khai, et autres
Publié: (2024)
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
par: Rai, Aashish, et autres
Publié: (2024)
par: Rai, Aashish, et autres
Publié: (2024)
LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading
par: Yemini, Yochai, et autres
Publié: (2023)
par: Yemini, Yochai, et autres
Publié: (2023)
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
par: Pham, Lam, et autres
Publié: (2024)
par: Pham, Lam, et autres
Publié: (2024)
Novel-View Acoustic Synthesis from 3D Reconstructed Rooms
par: Ahn, Byeongjoo, et autres
Publié: (2023)
par: Ahn, Byeongjoo, et autres
Publié: (2023)
The Impact of Frequency Bands on Acoustic Anomaly Detection of Machines using Deep Learning Based Model
par: Nguyen, Tin, et autres
Publié: (2024)
par: Nguyen, Tin, et autres
Publié: (2024)
MDSGen: Fast and Efficient Masked Diffusion Temporal-Aware Transformers for Open-Domain Sound Generation
par: Pham, Trung X., et autres
Publié: (2024)
par: Pham, Trung X., et autres
Publié: (2024)
Improving Acoustic Scene Classification with City Features
par: Cai, Yiqiang, et autres
Publié: (2025)
par: Cai, Yiqiang, et autres
Publié: (2025)
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
par: Choi, Jeongsoo, et autres
Publié: (2024)
par: Choi, Jeongsoo, et autres
Publié: (2024)
SOAF: Scene Occlusion-aware Neural Acoustic Field
par: Gao, Huiyu, et autres
Publié: (2024)
par: Gao, Huiyu, et autres
Publié: (2024)
Sonicmesh: Enhancing 3D Human Mesh Reconstruction in Vision-Impaired Environments With Acoustic Signals
par: Liang, Xiaoxuan, et autres
Publié: (2024)
par: Liang, Xiaoxuan, et autres
Publié: (2024)
UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation
par: Wang, Jinting, et autres
Publié: (2025)
par: Wang, Jinting, et autres
Publié: (2025)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
par: Chen, Wenxi, et autres
Publié: (2025)
par: Chen, Wenxi, et autres
Publié: (2025)
Improving Code Switching with Supervised Fine Tuning and GELU Adapters
par: Pham, Linh
Publié: (2025)
par: Pham, Linh
Publié: (2025)
VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos
par: Lin, Yan-Bo, et autres
Publié: (2024)
par: Lin, Yan-Bo, et autres
Publié: (2024)
Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis
par: Zhang, Zeyi, et autres
Publié: (2024)
par: Zhang, Zeyi, et autres
Publié: (2024)
Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing
par: Zhang, Zhedong, et autres
Publié: (2025)
par: Zhang, Zhedong, et autres
Publié: (2025)
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
par: Chen, Xueyuan, et autres
Publié: (2024)
par: Chen, Xueyuan, et autres
Publié: (2024)
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
par: Nguyen, Hong, et autres
Publié: (2024)
par: Nguyen, Hong, et autres
Publié: (2024)
Improving Streaming Speech Recognition With Time-Shifted Contextual Attention And Dynamic Right Context Masking
par: Le, Khanh, et autres
Publié: (2025)
par: Le, Khanh, et autres
Publié: (2025)
Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation
par: Li, Hao, et autres
Publié: (2025)
par: Li, Hao, et autres
Publié: (2025)
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
par: Xing, Jingyuan, et autres
Publié: (2025)
par: Xing, Jingyuan, et autres
Publié: (2025)
SAVE: Segment Audio-Visual Easy way using Segment Anything Model
par: Nguyen, Khanh-Binh, et autres
Publié: (2024)
par: Nguyen, Khanh-Binh, et autres
Publié: (2024)
PrismAudio: Decomposed Chain-of-Thoughts and Multi-dimensional Rewards for Video-to-Audio Generation
par: Liu, Huadai, et autres
Publié: (2025)
par: Liu, Huadai, et autres
Publié: (2025)
FilmComposer: LLM-Driven Music Production for Silent Film Clips
par: Xie, Zhifeng, et autres
Publié: (2025)
par: Xie, Zhifeng, et autres
Publié: (2025)
Sing-On-Your-Beat: Simple Text-Controllable Accompaniment Generations
par: Trinh, Quoc-Huy, et autres
Publié: (2024)
par: Trinh, Quoc-Huy, et autres
Publié: (2024)
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
par: Kim, Ji-Hoon, et autres
Publié: (2023)
par: Kim, Ji-Hoon, et autres
Publié: (2023)
DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
par: Zhang, Haomin, et autres
Publié: (2025)
par: Zhang, Haomin, et autres
Publié: (2025)
MuteSwap: Visual-informed Silent Video Identity Conversion
par: Liu, Yifan, et autres
Publié: (2025)
par: Liu, Yifan, et autres
Publié: (2025)
Few-shot Acoustic Synthesis with Multimodal Flow Matching
par: Brunetto, Amandine
Publié: (2026)
par: Brunetto, Amandine
Publié: (2026)
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
par: Wang, Haoran, et autres
Publié: (2025)
par: Wang, Haoran, et autres
Publié: (2025)
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
par: Hong, Joanna, et autres
Publié: (2025)
par: Hong, Joanna, et autres
Publié: (2025)
Documents similaires
-
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
par: Pham, The Hieu, et autres
Publié: (2025) -
Emotional Vietnamese Speech-Based Depression Diagnosis Using Dynamic Attention Mechanism
par: D., Quang-Anh N., et autres
Publié: (2024) -
Optimal Scalogram for Computational Complexity Reduction in Acoustic Recognition Using Deep Learning
par: Phan, Dang Thoai, et autres
Publié: (2025) -
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
par: Le, Khanh, et autres
Publié: (2025) -
Shushing! Let's Imagine an Authentic Speech from the Silent Video
par: Ye, Jiaxin, et autres
Publié: (2025)