T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Chung, Yoonjin, Lee, Junwon, Nam, Juhan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
di: Kim, Hounsu, et al.
Pubblicazione: (2024)
di: Kim, Hounsu, et al.
Pubblicazione: (2024)
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
di: Lee, Junwon, et al.
Pubblicazione: (2024)
di: Lee, Junwon, et al.
Pubblicazione: (2024)
CONMOD: Controllable Neural Frame-based Modulation Effects
di: Lee, Gyubin, et al.
Pubblicazione: (2024)
di: Lee, Gyubin, et al.
Pubblicazione: (2024)
CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation
di: Lee, Gyubin, et al.
Pubblicazione: (2026)
di: Lee, Gyubin, et al.
Pubblicazione: (2026)
KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation
di: Chung, Yoonjin, et al.
Pubblicazione: (2025)
di: Chung, Yoonjin, et al.
Pubblicazione: (2025)
Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
di: Lee, Jin Woo, et al.
Pubblicazione: (2024)
di: Lee, Jin Woo, et al.
Pubblicazione: (2024)
Wavetable Synthesis Using CVAE for Timbre Control Based on Semantic Label
di: Yutani, Tsugumasa, et al.
Pubblicazione: (2024)
di: Yutani, Tsugumasa, et al.
Pubblicazione: (2024)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2024)
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2024)
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
di: Cho, Hyunjae, et al.
Pubblicazione: (2024)
di: Cho, Hyunjae, et al.
Pubblicazione: (2024)
PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation
di: Lee, Sang-Hoon, et al.
Pubblicazione: (2024)
di: Lee, Sang-Hoon, et al.
Pubblicazione: (2024)
Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization
di: Lee, Sang-Hoon, et al.
Pubblicazione: (2024)
di: Lee, Sang-Hoon, et al.
Pubblicazione: (2024)
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
di: Ronchini, Francesca, et al.
Pubblicazione: (2025)
di: Ronchini, Francesca, et al.
Pubblicazione: (2025)
Sound Scene Synthesis at the DCASE 2024 Challenge
di: Lagrange, Mathieu, et al.
Pubblicazione: (2025)
di: Lagrange, Mathieu, et al.
Pubblicazione: (2025)
IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
di: Berger, Clémentine, et al.
Pubblicazione: (2025)
di: Berger, Clémentine, et al.
Pubblicazione: (2025)
Classification of Heart Sounds Using Multi-Branch Deep Convolutional Network and LSTM-CNN
di: Latifi, Seyed Amir, et al.
Pubblicazione: (2024)
di: Latifi, Seyed Amir, et al.
Pubblicazione: (2024)
A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music Modeling
di: Guo, Z., et al.
Pubblicazione: (2022)
di: Guo, Z., et al.
Pubblicazione: (2022)
EEG-Based Speech Decoding: A Novel Approach Using Multi-Kernel Ensemble Diffusion Models
di: Kim, Soowon, et al.
Pubblicazione: (2024)
di: Kim, Soowon, et al.
Pubblicazione: (2024)
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
di: Zhang, Yaoyun, et al.
Pubblicazione: (2024)
di: Zhang, Yaoyun, et al.
Pubblicazione: (2024)
CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical Controls
di: Chai, Li, et al.
Pubblicazione: (2024)
di: Chai, Li, et al.
Pubblicazione: (2024)
Dialogue in Resonance: An Interactive Music Piece for Piano and Real-Time Automatic Transcription System
di: Bang, Hayeon, et al.
Pubblicazione: (2025)
di: Bang, Hayeon, et al.
Pubblicazione: (2025)
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
di: Gao, Xiaoxue, et al.
Pubblicazione: (2025)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2025)
Pitch-Conditioned Instrument Sound Synthesis From an Interactive Timbre Latent Space
di: Limberg, Christian, et al.
Pubblicazione: (2025)
di: Limberg, Christian, et al.
Pubblicazione: (2025)
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
di: Berghi, Davide, et al.
Pubblicazione: (2025)
di: Berghi, Davide, et al.
Pubblicazione: (2025)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
di: Liu, Haohe, et al.
Pubblicazione: (2024)
di: Liu, Haohe, et al.
Pubblicazione: (2024)
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
di: Jung, Jihoo, et al.
Pubblicazione: (2026)
di: Jung, Jihoo, et al.
Pubblicazione: (2026)
UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching
di: Choi, Woongjib, et al.
Pubblicazione: (2025)
di: Choi, Woongjib, et al.
Pubblicazione: (2025)
FoleyBench: A Benchmark For Video-to-Audio Models
di: Dixit, Satvik, et al.
Pubblicazione: (2025)
di: Dixit, Satvik, et al.
Pubblicazione: (2025)
When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds
di: Kang, Minsu, et al.
Pubblicazione: (2025)
di: Kang, Minsu, et al.
Pubblicazione: (2025)
D3RM: A Discrete Denoising Diffusion Refinement Model for Piano Transcription
di: Kim, Hounsu, et al.
Pubblicazione: (2025)
di: Kim, Hounsu, et al.
Pubblicazione: (2025)
Learning Temporal Resolution in Spectrogram for Audio Classification
di: Liu, Haohe, et al.
Pubblicazione: (2022)
di: Liu, Haohe, et al.
Pubblicazione: (2022)
Speech Enhancement Based on Drifting Models
di: Xu, Liang, et al.
Pubblicazione: (2026)
di: Xu, Liang, et al.
Pubblicazione: (2026)
Evaluating the Temporal Detection Capability of Integrated Gradients Applied on Sound Classifier
di: Dumpis, Martynas, et al.
Pubblicazione: (2026)
di: Dumpis, Martynas, et al.
Pubblicazione: (2026)
A Hybrid Model for Weakly-Supervised Speech Dereverberation
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
DIFFRENT: A Diffusion Model for Recording Environment Transfer of Speech
di: Im, Jaekwon, et al.
Pubblicazione: (2024)
di: Im, Jaekwon, et al.
Pubblicazione: (2024)
U-DREAM: Unsupervised Dereverberation guided by a Reverberation Model
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
di: Bae, Hanbin, et al.
Pubblicazione: (2024)
di: Bae, Hanbin, et al.
Pubblicazione: (2024)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
di: Gállego, Gerard I., et al.
Pubblicazione: (2024)
di: Gállego, Gerard I., et al.
Pubblicazione: (2024)
Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
di: Huang, Kuan-Tang, et al.
Pubblicazione: (2026)
di: Huang, Kuan-Tang, et al.
Pubblicazione: (2026)
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024)
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
di: Kim, Hounsu, et al.
Pubblicazione: (2024) -
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
di: Lee, Junwon, et al.
Pubblicazione: (2024) -
CONMOD: Controllable Neural Frame-based Modulation Effects
di: Lee, Gyubin, et al.
Pubblicazione: (2024) -
CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation
di: Lee, Gyubin, et al.
Pubblicazione: (2026) -
KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation
di: Chung, Yoonjin, et al.
Pubblicazione: (2025)