Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Chung, Soo-Whan, Choi, Min-Seok |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
by: Lee, Yongjoon, et al.
Published: (2026)
by: Lee, Yongjoon, et al.
Published: (2026)
Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation
by: Kim, Miseul, et al.
Published: (2024)
by: Kim, Miseul, et al.
Published: (2024)
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
by: Cha, Jun-Hyeok, et al.
Published: (2025)
by: Cha, Jun-Hyeok, et al.
Published: (2025)
Reproducing the Acoustic Velocity Vectors in a Circular Listening Area
by: Wang, Jiarui, et al.
Published: (2024)
by: Wang, Jiarui, et al.
Published: (2024)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
by: Cho, Deok-Hyeon, et al.
Published: (2025)
by: Cho, Deok-Hyeon, et al.
Published: (2025)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
by: Oh, Hyung-Seok, et al.
Published: (2023)
by: Oh, Hyung-Seok, et al.
Published: (2023)
Acoustic BPE for Speech Generation with Discrete Tokens
by: Shen, Feiyu, et al.
Published: (2023)
by: Shen, Feiyu, et al.
Published: (2023)
DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
by: Oh, Hyung-Seok, et al.
Published: (2024)
by: Oh, Hyung-Seok, et al.
Published: (2024)
Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
by: Cai, Yiqiang, et al.
Published: (2024)
by: Cai, Yiqiang, et al.
Published: (2024)
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
by: Jeong, Jihoon, et al.
Published: (2026)
by: Jeong, Jihoon, et al.
Published: (2026)
Evaluating Speech Enhancement Systems Through Listening Effort
by: Gelderblom, Femke B., et al.
Published: (2024)
by: Gelderblom, Femke B., et al.
Published: (2024)
SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features
by: Shi, Yu-Fei, et al.
Published: (2024)
by: Shi, Yu-Fei, et al.
Published: (2024)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
by: Ku, Pin-Jui, et al.
Published: (2024)
by: Ku, Pin-Jui, et al.
Published: (2024)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
by: Cho, Deok-Hyeon, et al.
Published: (2025)
by: Cho, Deok-Hyeon, et al.
Published: (2025)
Comparative Evaluation of Acoustic Feature Extraction Tools for Clinical Speech Analysis
by: Choi, Anna Seo Gyeong, et al.
Published: (2025)
by: Choi, Anna Seo Gyeong, et al.
Published: (2025)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
by: Chen, Zhengyang, et al.
Published: (2024)
by: Chen, Zhengyang, et al.
Published: (2024)
Sound Field Synthesis with Acoustic Waves
by: Mansour, Mohamed F.
Published: (2024)
by: Mansour, Mohamed F.
Published: (2024)
Listen and Move: Improving GANs Coherency in Agnostic Sound-to-Video Generation
by: Redondo, Rafael
Published: (2024)
by: Redondo, Rafael
Published: (2024)
Leveraging Sound Source Trajectories for Universal Sound Separation
by: Wu, Donghang, et al.
Published: (2024)
by: Wu, Donghang, et al.
Published: (2024)
FLOWER: Flow-Based Estimated Gaussian Guidance for General Speech Restoration
by: Yang, Da-Hee, et al.
Published: (2025)
by: Yang, Da-Hee, et al.
Published: (2025)
ReverbMiipher: Generative Speech Restoration meets Reverberation Characteristics Controllability
by: Nakata, Wataru, et al.
Published: (2025)
by: Nakata, Wataru, et al.
Published: (2025)
Learning How to Listen: A Temporal-Frequential Attention Model for Sound Event Detection
by: Shen, Yu-Han, et al.
Published: (2018)
by: Shen, Yu-Han, et al.
Published: (2018)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
by: Yen, Hao, et al.
Published: (2024)
by: Yen, Hao, et al.
Published: (2024)
A Steered Response Power Method for Sound Source Localization With Generic Acoustic Models
by: Müller, Kaspar, et al.
Published: (2025)
by: Müller, Kaspar, et al.
Published: (2025)
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
by: Kirdey, Stanislav
Published: (2025)
by: Kirdey, Stanislav
Published: (2025)
SoundCompass: Navigating Target Sound Extraction With Effective Directional Clue Integration In Complex Acoustic Scenes
by: Choi, Dayun, et al.
Published: (2025)
by: Choi, Dayun, et al.
Published: (2025)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
by: Hu, Cheng-Hung, et al.
Published: (2025)
by: Hu, Cheng-Hung, et al.
Published: (2025)
VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion
by: Choi, Joon-Seung, et al.
Published: (2025)
by: Choi, Joon-Seung, et al.
Published: (2025)
TF-CorrNet: Leveraging Spatial Correlation for Continuous Speech Separation
by: Shin, Ui-Hyeop, et al.
Published: (2025)
by: Shin, Ui-Hyeop, et al.
Published: (2025)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
by: Saeki, Takaaki, et al.
Published: (2024)
by: Saeki, Takaaki, et al.
Published: (2024)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
by: Nam, KiHyun, et al.
Published: (2024)
by: Nam, KiHyun, et al.
Published: (2024)
Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners
by: Yamamoto, Katsuhiko, et al.
Published: (2025)
by: Yamamoto, Katsuhiko, et al.
Published: (2025)
Leveraging Multiple Speech Enhancers for Non-Intrusive Intelligibility Prediction for Hearing-Impaired Listeners
by: Cao, Boxuan, et al.
Published: (2025)
by: Cao, Boxuan, et al.
Published: (2025)
Retaining Mixture Representations for Domain Generalized Anomalous Sound Detection
by: Saengthong, Phurich, et al.
Published: (2025)
by: Saengthong, Phurich, et al.
Published: (2025)
Benchmarking Representations for Speech, Music, and Acoustic Events
by: La Quatra, Moreno, et al.
Published: (2024)
by: La Quatra, Moreno, et al.
Published: (2024)
LL-SDR: Low-Latency Speech enhancement through Discrete Representations
by: Li, Jingyi, et al.
Published: (2026)
by: Li, Jingyi, et al.
Published: (2026)
Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes
by: Nam, Hyeonuk, et al.
Published: (2024)
by: Nam, Hyeonuk, et al.
Published: (2024)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
by: Meghanani, Amit, et al.
Published: (2024)
by: Meghanani, Amit, et al.
Published: (2024)
Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis
by: Chung, Raymond
Published: (2026)
by: Chung, Raymond
Published: (2026)
RF-GML: Reference-Free Generative Machine Listener
by: Biswas, Arijit, et al.
Published: (2024)
by: Biswas, Arijit, et al.
Published: (2024)
Similar Items
-
SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
by: Lee, Yongjoon, et al.
Published: (2026) -
Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation
by: Kim, Miseul, et al.
Published: (2024) -
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
by: Cha, Jun-Hyeok, et al.
Published: (2025) -
Reproducing the Acoustic Velocity Vectors in a Circular Listening Area
by: Wang, Jiarui, et al.
Published: (2024) -
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
by: Cho, Deok-Hyeon, et al.
Published: (2025)