Recomposer: Event-roll-guided generative audio editing
Fuente:
arXiv
Saved in:
| Main Authors: | Ellis, Daniel P. W., Fonseca, Eduardo, Weiss, Ron J., Wilson, Kevin, Wisdom, Scott, Erdogan, Hakan, Hershey, John R., Jansen, Aren, Moore, R. Channing, Plakal, Manoj |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsupervised Multi-channel Separation and Adaptation
by: Han, Cong, et al.
Published: (2023)
by: Han, Cong, et al.
Published: (2023)
Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables
by: Dementyev, Artem, et al.
Published: (2024)
by: Dementyev, Artem, et al.
Published: (2024)
The CHiME-7 UDASE task: Unsupervised domain adaptation for conversational speech enhancement
by: Leglaive, Simon, et al.
Published: (2023)
by: Leglaive, Simon, et al.
Published: (2023)
AudioMorphix: Training-free audio editing with diffusion probabilistic models
by: Liang, Jinhua, et al.
Published: (2025)
by: Liang, Jinhua, et al.
Published: (2025)
SCORE: Scaling audio generation using Standardized COmposite REwards
by: Jung, Jaemin, et al.
Published: (2025)
by: Jung, Jaemin, et al.
Published: (2025)
Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge
by: Leglaive, Simon, et al.
Published: (2024)
by: Leglaive, Simon, et al.
Published: (2024)
SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
by: Chen, Tuochao, et al.
Published: (2025)
by: Chen, Tuochao, et al.
Published: (2025)
Source Separation by Flow Matching
by: Scheibler, Robin, et al.
Published: (2025)
by: Scheibler, Robin, et al.
Published: (2025)
STASE: A spatialized text-to-audio synthesis engine for music generation
by: Chi, Tutti, et al.
Published: (2025)
by: Chi, Tutti, et al.
Published: (2025)
DashengTokenizer: One layer is enough for unified audio understanding and generation
by: Dinkel, Heinrich, et al.
Published: (2026)
by: Dinkel, Heinrich, et al.
Published: (2026)
ImmersiveFlow: Stereo-to-7.1.4 spatial audio generation with flow matching
by: Liang, Zining, et al.
Published: (2026)
by: Liang, Zining, et al.
Published: (2026)
A SOUND APPROACH: Using Large Language Models to generate audio descriptions for egocentric text-audio retrieval
by: Oncescu, Andreea-Maria, et al.
Published: (2024)
by: Oncescu, Andreea-Maria, et al.
Published: (2024)
Unsupervised Improved MVDR Beamforming for Sound Enhancement
by: Kealey, Jacob, et al.
Published: (2024)
by: Kealey, Jacob, et al.
Published: (2024)
Visual-based spatial audio generation system for multi-speaker environments
by: Liu, Xiaojing, et al.
Published: (2025)
by: Liu, Xiaojing, et al.
Published: (2025)
Online incremental learning for audio classification using a pretrained audio model
by: Mulimani, Manjunath, et al.
Published: (2025)
by: Mulimani, Manjunath, et al.
Published: (2025)
Testing chatbots on the creation of encoders for audio conditioned image generation
by: León, Jorge E., et al.
Published: (2025)
by: León, Jorge E., et al.
Published: (2025)
Scaling up masked audio encoder learning for general audio classification
by: Dinkel, Heinrich, et al.
Published: (2024)
by: Dinkel, Heinrich, et al.
Published: (2024)
A tunable binaural audio telepresence system capable of balancing immersive and enhanced modes
by: Hsu, Yicheng, et al.
Published: (2024)
by: Hsu, Yicheng, et al.
Published: (2024)
AEROMamba: An efficient architecture for audio super-resolution using generative adversarial networks and state space models
by: Abreu, Wallace, et al.
Published: (2024)
by: Abreu, Wallace, et al.
Published: (2024)
Enhanced Sound Event Localization and Detection in Real 360-degree audio-visual soundscapes
by: Roman, Adrian S., et al.
Published: (2024)
by: Roman, Adrian S., et al.
Published: (2024)
Synthetic training set generation using text-to-audio models for environmental sound classification
by: Ronchini, Francesca, et al.
Published: (2024)
by: Ronchini, Francesca, et al.
Published: (2024)
Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio question answering
by: Zhao, Jinghua, et al.
Published: (2025)
by: Zhao, Jinghua, et al.
Published: (2025)
Long-Form Speech Generation with Spoken Language Models
by: Park, Se Jin, et al.
Published: (2024)
by: Park, Se Jin, et al.
Published: (2024)
Towards audio language modeling -- an overview
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
WavLM model ensemble for audio deepfake detection
by: Combei, David, et al.
Published: (2024)
by: Combei, David, et al.
Published: (2024)
Cryfish: On deep audio analysis with Large Language Models
by: Mitrofanov, Anton, et al.
Published: (2025)
by: Mitrofanov, Anton, et al.
Published: (2025)
Multiple Hankel matrix rank minimization for audio inpainting
by: Záviška, Pavel, et al.
Published: (2023)
by: Záviška, Pavel, et al.
Published: (2023)
Where are we in audio deepfake detection? A systematic analysis over generative and detection models
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Binaural Angular Separation Network
by: Yang, Yang, et al.
Published: (2024)
by: Yang, Yang, et al.
Published: (2024)
Are audio DeepFake detection models polyglots?
by: Marek, Bartłomiej, et al.
Published: (2024)
by: Marek, Bartłomiej, et al.
Published: (2024)
Tweaking autoregressive methods for inpainting of gaps in audio signals
by: Mokrý, Ondřej, et al.
Published: (2024)
by: Mokrý, Ondřej, et al.
Published: (2024)
MBCodec:Thorough disentangle for high-fidelity audio compression
by: Zhang, Ruonan, et al.
Published: (2025)
by: Zhang, Ruonan, et al.
Published: (2025)
Real-time implementation of vibrato transfer as an audio effect
by: Hyrkas, Jeremy
Published: (2025)
by: Hyrkas, Jeremy
Published: (2025)
Towards predicting binaural audio quality in listeners with normal and impaired hearing
by: Biberger, Thomas, et al.
Published: (2025)
by: Biberger, Thomas, et al.
Published: (2025)
Sound event detection with audio-text models and heterogeneous temporal annotations
by: Harju, Manu, et al.
Published: (2025)
by: Harju, Manu, et al.
Published: (2025)
Multi-label audio classification with a noisy zero-shot teacher
by: Braun, Sebastian, et al.
Published: (2024)
by: Braun, Sebastian, et al.
Published: (2024)
Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
by: Pîrlogeanu, Gabriel, et al.
Published: (2026)
by: Pîrlogeanu, Gabriel, et al.
Published: (2026)
Unmasking real-world audio deepfakes: A data-centric approach
by: Combei, David, et al.
Published: (2025)
by: Combei, David, et al.
Published: (2025)
LVNS-RAVE: Diversified audio generation with RAVE and Latent Vector Novelty Search
by: Guo, Jinyue, et al.
Published: (2024)
by: Guo, Jinyue, et al.
Published: (2024)
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
by: Rouard, Simon, et al.
Published: (2025)
by: Rouard, Simon, et al.
Published: (2025)
Similar Items
-
Unsupervised Multi-channel Separation and Adaptation
by: Han, Cong, et al.
Published: (2023) -
Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables
by: Dementyev, Artem, et al.
Published: (2024) -
The CHiME-7 UDASE task: Unsupervised domain adaptation for conversational speech enhancement
by: Leglaive, Simon, et al.
Published: (2023) -
AudioMorphix: Training-free audio editing with diffusion probabilistic models
by: Liang, Jinhua, et al.
Published: (2025) -
SCORE: Scaling audio generation using Standardized COmposite REwards
by: Jung, Jaemin, et al.
Published: (2025)