Recomposer: Event-roll-guided generative audio editing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914024387510272 |
|---|---|
| author | Ellis, Daniel P. W. Fonseca, Eduardo Weiss, Ron J. Wilson, Kevin Wisdom, Scott Erdogan, Hakan Hershey, John R. Jansen, Aren Moore, R. Channing Plakal, Manoj |
| author_facet | Ellis, Daniel P. W. Fonseca, Eduardo Weiss, Ron J. Wilson, Kevin Wisdom, Scott Erdogan, Hakan Hershey, John R. Jansen, Aren Moore, R. Channing Plakal, Manoj |
| contents | Editing complex real-world sound scenes is difficult because individual sound sources overlap in time. Generative models can fill-in missing or corrupted details based on their strong prior understanding of the data domain. We present a system for editing individual sound events within complex scenes able to delete, insert, and enhance individual sound events based on textual edit descriptions (e.g., ``enhance Door'') and a graphical representation of the event timing derived from an ``event roll'' transcription. We present an encoder-decoder transformer working on SoundStream representations, trained on synthetic (input, desired output) audio example pairs formed by adding isolated sound events to dense, real-world backgrounds. Evaluation reveals the importance of each part of the edit descriptions -- action, class, timing. Our work demonstrates ``recomposition'' is an important and practical application. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_05256 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Recomposer: Event-roll-guided generative audio editing Ellis, Daniel P. W. Fonseca, Eduardo Weiss, Ron J. Wilson, Kevin Wisdom, Scott Erdogan, Hakan Hershey, John R. Jansen, Aren Moore, R. Channing Plakal, Manoj Sound Artificial Intelligence Machine Learning Audio and Speech Processing Editing complex real-world sound scenes is difficult because individual sound sources overlap in time. Generative models can fill-in missing or corrupted details based on their strong prior understanding of the data domain. We present a system for editing individual sound events within complex scenes able to delete, insert, and enhance individual sound events based on textual edit descriptions (e.g., ``enhance Door'') and a graphical representation of the event timing derived from an ``event roll'' transcription. We present an encoder-decoder transformer working on SoundStream representations, trained on synthetic (input, desired output) audio example pairs formed by adding isolated sound events to dense, real-world backgrounds. Evaluation reveals the importance of each part of the edit descriptions -- action, class, timing. Our work demonstrates ``recomposition'' is an important and practical application. |
| title | Recomposer: Event-roll-guided generative audio editing |
| topic | Sound Artificial Intelligence Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2509.05256 |