SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Zhi, Takahashi, Akira, Cui, Shuyang, Toyama, Keisuke, Takahashi, Shusuke, Mitsufuji, Yuki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
by: Comunità, Marco, et al.
Published: (2024)
by: Comunità, Marco, et al.
Published: (2024)
Do Foundational Audio Encoders Understand Music Structure?
by: Toyama, Keisuke, et al.
Published: (2025)
by: Toyama, Keisuke, et al.
Published: (2025)
The Whole Is Greater than the Sum of Its Parts: Improving Music Source Separation by Bridging Network
by: Sawata, Ryosuke, et al.
Published: (2023)
by: Sawata, Ryosuke, et al.
Published: (2023)
MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
by: Takahashi, Akira, et al.
Published: (2025)
by: Takahashi, Akira, et al.
Published: (2025)
Diffusion-based Signal Refiner for Speech Enhancement and Separation
by: Hirano, Masato, et al.
Published: (2023)
by: Hirano, Masato, et al.
Published: (2023)
Schrödinger Bridge Consistency Trajectory Models for Speech Enhancement
by: Nishigori, Shuichiro, et al.
Published: (2025)
by: Nishigori, Shuichiro, et al.
Published: (2025)
MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation
by: Takahashi, Akira, et al.
Published: (2026)
by: Takahashi, Akira, et al.
Published: (2026)
SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation
by: Shimada, Kazuki, et al.
Published: (2024)
by: Shimada, Kazuki, et al.
Published: (2024)
Noise-to-Notes: Diffusion-based Generation and Refinement for Automatic Drum Transcription
by: Yeung, Michael, et al.
Published: (2025)
by: Yeung, Michael, et al.
Published: (2025)
Zero- and Few-shot Sound Event Localization and Detection
by: Shimada, Kazuki, et al.
Published: (2023)
by: Shimada, Kazuki, et al.
Published: (2023)
FoleyBench: A Benchmark For Video-to-Audio Models
by: Dixit, Satvik, et al.
Published: (2025)
by: Dixit, Satvik, et al.
Published: (2025)
DiffRoll: Diffusion-based Generative Music Transcription with Unsupervised Pretraining Capability
by: Cheuk, Kin Wai, et al.
Published: (2022)
by: Cheuk, Kin Wai, et al.
Published: (2022)
Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
by: Shi, Hao, et al.
Published: (2023)
by: Shi, Hao, et al.
Published: (2023)
Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
by: Hou, Siyuan, et al.
Published: (2024)
by: Hou, Siyuan, et al.
Published: (2024)
TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
by: Jeong, Jaeseok, et al.
Published: (2025)
by: Jeong, Jaeseok, et al.
Published: (2025)
MambaFoley: Foley Sound Generation using Selective State-Space Models
by: Colombo, Marco Furio, et al.
Published: (2024)
by: Colombo, Marco Furio, et al.
Published: (2024)
SilentCipher: Deep Audio Watermarking
by: Singh, Mayank Kumar, et al.
Published: (2024)
by: Singh, Mayank Kumar, et al.
Published: (2024)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
by: Takano, Taisei, et al.
Published: (2025)
by: Takano, Taisei, et al.
Published: (2025)
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
by: Kanamori, Yusuke, et al.
Published: (2025)
by: Kanamori, Yusuke, et al.
Published: (2025)
Spectral Masking with Explicit Time-Context Windowing for Neural Network-Based Monaural Speech Enhancement
by: Fiorio, Luan Vinícius, et al.
Published: (2024)
by: Fiorio, Luan Vinícius, et al.
Published: (2024)
CAFA: a Controllable Automatic Foley Artist
by: Benita, Roi, et al.
Published: (2025)
by: Benita, Roi, et al.
Published: (2025)
Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music Transcription
by: Cwitkowitz, Frank, et al.
Published: (2023)
by: Cwitkowitz, Frank, et al.
Published: (2023)
The Sound Demixing Challenge 2023 $\unicode{x2013}$ Cinematic Demixing Track
by: Uhlich, Stefan, et al.
Published: (2023)
by: Uhlich, Stefan, et al.
Published: (2023)
MaskBeat: Loopable Drum Beat Generation
by: Lanzendörfer, Luca A., et al.
Published: (2025)
by: Lanzendörfer, Luca A., et al.
Published: (2025)
Masked Audio Modeling with CLAP and Multi-Objective Learning
by: Xin, Yifei, et al.
Published: (2024)
by: Xin, Yifei, et al.
Published: (2024)
Scaling up masked audio encoder learning for general audio classification
by: Dinkel, Heinrich, et al.
Published: (2024)
by: Dinkel, Heinrich, et al.
Published: (2024)
OpenMU: Your Swiss Army Knife for Music Understanding
by: Zhao, Mengjie, et al.
Published: (2024)
by: Zhao, Mengjie, et al.
Published: (2024)
Ambisonics Binaural Rendering via Masked Magnitude Least Squares
by: Berebi, Or, et al.
Published: (2025)
by: Berebi, Or, et al.
Published: (2025)
Disentangling Dual-Encoder Masked Autoencoder for Respiratory Sound Classification
by: Wei, Peidong, et al.
Published: (2025)
by: Wei, Peidong, et al.
Published: (2025)
LiLAC: A Lightweight Latent ControlNet for Musical Audio Generation
by: Baker, Tom, et al.
Published: (2025)
by: Baker, Tom, et al.
Published: (2025)
Audio Palette: A Diffusion Transformer with Multi-Signal Conditioning for Controllable Foley Synthesis
by: Wang, Junnuo
Published: (2025)
by: Wang, Junnuo
Published: (2025)
Supervised contrastive learning from weakly-labeled audio segments for musical version matching
by: Serrà, Joan, et al.
Published: (2025)
by: Serrà, Joan, et al.
Published: (2025)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
by: Le, Khanh, et al.
Published: (2025)
by: Le, Khanh, et al.
Published: (2025)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
by: Chen, Li-Wei, et al.
Published: (2024)
by: Chen, Li-Wei, et al.
Published: (2024)
The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
by: O'Reilly, Patrick, et al.
Published: (2025)
by: O'Reilly, Patrick, et al.
Published: (2025)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
by: Pham, The Hieu, et al.
Published: (2025)
by: Pham, The Hieu, et al.
Published: (2025)
Trainingless Adaptation of Pretrained Models for Environmental Sound Classification
by: Tonami, Noriyuki, et al.
Published: (2024)
by: Tonami, Noriyuki, et al.
Published: (2024)
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
EEND-M2F: Masked-attention mask transformers for speaker diarization
by: Härkönen, Marc, et al.
Published: (2024)
by: Härkönen, Marc, et al.
Published: (2024)
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
by: Xi, Yu, et al.
Published: (2025)
by: Xi, Yu, et al.
Published: (2025)
Similar Items
-
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
by: Comunità, Marco, et al.
Published: (2024) -
Do Foundational Audio Encoders Understand Music Structure?
by: Toyama, Keisuke, et al.
Published: (2025) -
The Whole Is Greater than the Sum of Its Parts: Improving Music Source Separation by Bridging Network
by: Sawata, Ryosuke, et al.
Published: (2023) -
MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
by: Takahashi, Akira, et al.
Published: (2025) -
Diffusion-based Signal Refiner for Speech Enhancement and Separation
by: Hirano, Masato, et al.
Published: (2023)