Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Miseul, Chung, Soo-Whan, Ji, Youna, Kang, Hong-Goo, Choi, Min-Seok |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpeechMLC: Speech Multi-label Classification
by: Kim, Miseul, et al.
Published: (2025)
by: Kim, Miseul, et al.
Published: (2025)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
by: Chung, Soo-Whan, et al.
Published: (2025)
by: Chung, Soo-Whan, et al.
Published: (2025)
Speaker-Independent Acoustic-to-Articulatory Inversion through Multi-Channel Attention Discriminator
by: Chung, Woo-Jin, et al.
Published: (2024)
by: Chung, Woo-Jin, et al.
Published: (2024)
Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
by: Kim, Miseul, et al.
Published: (2025)
by: Kim, Miseul, et al.
Published: (2025)
MF-PAM: Accurate Pitch Estimation through Periodicity Analysis and Multi-level Feature Fusion
by: Chung, Woo-Jin, et al.
Published: (2023)
by: Chung, Woo-Jin, et al.
Published: (2023)
Differentiable Acoustic Radiance Transfer
by: Lee, Sungho, et al.
Published: (2025)
by: Lee, Sungho, et al.
Published: (2025)
Neural Spectral Band Generation for Audio Coding
by: Choi, Woongjib, et al.
Published: (2025)
by: Choi, Woongjib, et al.
Published: (2025)
A Speech Production Model for Radar: Connecting Speech Acoustics with Radar-Measured Vibrations
by: Lenz, Isabella, et al.
Published: (2025)
by: Lenz, Isabella, et al.
Published: (2025)
Analyzing the Impact of Accent on English Speech: Acoustic and Articulatory Perspectives
by: Premananth, Gowtham, et al.
Published: (2025)
by: Premananth, Gowtham, et al.
Published: (2025)
UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching
by: Choi, Woongjib, et al.
Published: (2025)
by: Choi, Woongjib, et al.
Published: (2025)
Lightweight and Generalizable Acoustic Scene Representations via Contrastive Fine-Tuning and Distillation
by: Yuan, Kuang, et al.
Published: (2025)
by: Yuan, Kuang, et al.
Published: (2025)
DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
by: Lee, Dongheon, et al.
Published: (2023)
by: Lee, Dongheon, et al.
Published: (2023)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
by: Kim, Ji-Hoon, et al.
Published: (2024)
by: Kim, Ji-Hoon, et al.
Published: (2024)
IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
by: Berger, Clémentine, et al.
Published: (2025)
by: Berger, Clémentine, et al.
Published: (2025)
Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation
by: Kim, Minsu, et al.
Published: (2023)
by: Kim, Minsu, et al.
Published: (2023)
Unsupervised Variational Acoustic Clustering
by: Fiorio, Luan Vinícius, et al.
Published: (2025)
by: Fiorio, Luan Vinícius, et al.
Published: (2025)
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
by: Nguyen, Tan Dat, et al.
Published: (2024)
by: Nguyen, Tan Dat, et al.
Published: (2024)
Speech Enhancement based on cascaded two flows
by: Lee, Seonggyu, et al.
Published: (2025)
by: Lee, Seonggyu, et al.
Published: (2025)
EchoScan: Scanning Complex Room Geometries via Acoustic Echoes
by: Yeon, Inmo, et al.
Published: (2023)
by: Yeon, Inmo, et al.
Published: (2023)
Cyclic Multichannel Wiener Filter for Acoustic Beamforming
by: Bologni, Giovanni, et al.
Published: (2025)
by: Bologni, Giovanni, et al.
Published: (2025)
Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
by: Kwon, Younghoo, et al.
Published: (2024)
by: Kwon, Younghoo, et al.
Published: (2024)
Transferable Selective Virtual Sensing Active Noise Control Technique Based on Metric Learning
by: Wang, Boxiang, et al.
Published: (2024)
by: Wang, Boxiang, et al.
Published: (2024)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
by: Yu, Chin-Yun, et al.
Published: (2022)
by: Yu, Chin-Yun, et al.
Published: (2022)
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
by: Yang, Shu-wen, et al.
Published: (2025)
by: Yang, Shu-wen, et al.
Published: (2025)
FUN-SSL: Full-band Layer Followed by U-Net with Narrow-band Layers for Multiple Moving Sound Source Localization
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
Acoustic Non-Stationarity Objective Assessment with Hard Label Criteria for Supervised Learning Models
by: Zucatelli, Guilherme, et al.
Published: (2025)
by: Zucatelli, Guilherme, et al.
Published: (2025)
Reduce Computational Complexity for Continuous Wavelet Transform in Acoustic Recognition Using Hop Size
by: Phan, Dang Thoai
Published: (2024)
by: Phan, Dang Thoai
Published: (2024)
Binaural Localization Model for Speech in Noise
by: Tokala, Vikas, et al.
Published: (2025)
by: Tokala, Vikas, et al.
Published: (2025)
Speech-Based Prioritization for Schizophrenia Intervention
by: Premananth, Gowtham, et al.
Published: (2025)
by: Premananth, Gowtham, et al.
Published: (2025)
Prompt-driven Target Speech Diarization
by: Jiang, Yidi, et al.
Published: (2023)
by: Jiang, Yidi, et al.
Published: (2023)
Frequency-Modulated and Single-Tone Excitation to Reveal Vibro-Acoustic Nonlinearities in Loosened Bolted Joints
by: Kullukcu, Berkay, et al.
Published: (2026)
by: Kullukcu, Berkay, et al.
Published: (2026)
Graph-Enhanced Dual-Stream Feature Fusion with Pre-Trained Model for Acoustic Traffic Monitoring
by: Fan, Shitong, et al.
Published: (2024)
by: Fan, Shitong, et al.
Published: (2024)
Brain-Informed Speech Separation for Cochlear Implants
by: Gajecki, Tom, et al.
Published: (2026)
by: Gajecki, Tom, et al.
Published: (2026)
Physics-Informed Neural Network-Driven Sparse Field Discretization Method for Near-Field Acoustic Holography
by: Luan, Xinmeng, et al.
Published: (2025)
by: Luan, Xinmeng, et al.
Published: (2025)
Ultrasensitive Textile Strain Sensors Redefine Wearable Silent Speech Interfaces with High Machine Learning Efficiency
by: Tang, Chenyu, et al.
Published: (2023)
by: Tang, Chenyu, et al.
Published: (2023)
A Study on Speech Assessment with Visual Cues
by: Ahmed, Shafique, et al.
Published: (2025)
by: Ahmed, Shafique, et al.
Published: (2025)
EMOCONV-DIFF: Diffusion-based Speech Emotion Conversion for Non-parallel and In-the-wild Data
by: Prabhu, Navin Raj, et al.
Published: (2023)
by: Prabhu, Navin Raj, et al.
Published: (2023)
FlowSE: Flow Matching-based Speech Enhancement
by: Lee, Seonggyu, et al.
Published: (2025)
by: Lee, Seonggyu, et al.
Published: (2025)
Enhanced Deep Speech Separation in Clustered Ad Hoc Distributed Microphone Environments
by: Kim, Jihyun, et al.
Published: (2024)
by: Kim, Jihyun, et al.
Published: (2024)
One-Shot Distributed Node-Specific Signal Estimation with Non-Overlapping Latent Subspaces in Acoustic Sensor Networks
by: Didier, Paul, et al.
Published: (2024)
by: Didier, Paul, et al.
Published: (2024)
Similar Items
-
SpeechMLC: Speech Multi-label Classification
by: Kim, Miseul, et al.
Published: (2025) -
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
by: Chung, Soo-Whan, et al.
Published: (2025) -
Speaker-Independent Acoustic-to-Articulatory Inversion through Multi-Channel Attention Discriminator
by: Chung, Woo-Jin, et al.
Published: (2024) -
Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
by: Kim, Miseul, et al.
Published: (2025) -
MF-PAM: Accurate Pitch Estimation through Periodicity Analysis and Multi-level Feature Fusion
by: Chung, Woo-Jin, et al.
Published: (2023)