Saved in:
| Main Authors: | Chu, Annie, García, Hugo Flores, Nieto, Oriol, Salamon, Justin, Pardo, Bryan, Seetharaman, Prem |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.20426 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative Audio Extension and Morphing
by: Seetharaman, Prem, et al.
Published: (2026)
by: Seetharaman, Prem, et al.
Published: (2026)
Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations
by: García, Hugo Flores, et al.
Published: (2024)
by: García, Hugo Flores, et al.
Published: (2024)
Audiocards: Structured Metadata Improves Audio Language Models For Sound Design
by: Sridhar, Sripathi, et al.
Published: (2026)
by: Sridhar, Sripathi, et al.
Published: (2026)
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
by: Kumar, Sonal, et al.
Published: (2024)
by: Kumar, Sonal, et al.
Published: (2024)
Video-Guided Foley Sound Generation with Multimodal Controls
by: Chen, Ziyang, et al.
Published: (2024)
by: Chen, Ziyang, et al.
Published: (2024)
AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing
by: Chen, William, et al.
Published: (2026)
by: Chen, William, et al.
Published: (2026)
The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
by: O'Reilly, Patrick, et al.
Published: (2025)
by: O'Reilly, Patrick, et al.
Published: (2025)
FLAM: Frame-Wise Language-Audio Modeling
by: Wu, Yusong, et al.
Published: (2025)
by: Wu, Yusong, et al.
Published: (2025)
TAC: Timestamped Audio Captioning
by: Kumar, Sonal, et al.
Published: (2026)
by: Kumar, Sonal, et al.
Published: (2026)
Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning
by: Manco, Ilaria, et al.
Published: (2024)
by: Manco, Ilaria, et al.
Published: (2024)
PromptSep: Generative Audio Separation via Multimodal Prompting
by: Wen, Yutong, et al.
Published: (2025)
by: Wen, Yutong, et al.
Published: (2025)
Code Drift: Towards Idempotent Neural Audio Codecs
by: O'Reilly, Patrick, et al.
Published: (2024)
by: O'Reilly, Patrick, et al.
Published: (2024)
SoundMorpher: Perceptually-Uniform Sound Morphing with Diffusion Model
by: Niu, Xinlei, et al.
Published: (2024)
by: Niu, Xinlei, et al.
Published: (2024)
Sound Effects and the Morphing Façade: Emotion in Sound of Projection Mapping
by: Roger Pastó Cortina
Published: (2019)
by: Roger Pastó Cortina
Published: (2019)
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
by: Chu, Annie, et al.
Published: (2024)
by: Chu, Annie, et al.
Published: (2024)
Time-Domain Voice Identity Morphing (TD-VIM): A Signal-Level Approach to Morphing Attacks on Speaker Verification Systems
by: PN, Aravinda Reddy, et al.
Published: (2026)
by: PN, Aravinda Reddy, et al.
Published: (2026)
Taming Audio VAEs via Target-KL Regularization
by: Seetharaman, Prem, et al.
Published: (2026)
by: Seetharaman, Prem, et al.
Published: (2026)
Exploring Musical Roots: Applying Audio Embeddings to Empower Influence Attribution for a Generative Music Model
by: Barnett, Julia, et al.
Published: (2024)
by: Barnett, Julia, et al.
Published: (2024)
Learning Perceptually Relevant Temporal Envelope Morphing
by: Dixit, Satvik, et al.
Published: (2025)
by: Dixit, Satvik, et al.
Published: (2025)
VoxMorph: Scalable Zero-shot Voice Identity Morphing via Disentangled Embeddings
by: Krishnamurthy, Bharath, et al.
Published: (2026)
by: Krishnamurthy, Bharath, et al.
Published: (2026)
Towards Controllable Audio Texture Morphing
by: Gupta, Chitralekha, et al.
Published: (2023)
by: Gupta, Chitralekha, et al.
Published: (2023)
Ethics Statements in AI Music Papers: The Effective and the Ineffective
by: Barnett, Julia, et al.
Published: (2025)
by: Barnett, Julia, et al.
Published: (2025)
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
SoundShift: Exploring Sound Manipulations for Accessible Mixed-Reality Awareness
by: Chang, Ruei-Che, et al.
Published: (2024)
by: Chang, Ruei-Che, et al.
Published: (2024)
Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
by: Bae, Sangmin, et al.
Published: (2023)
by: Bae, Sangmin, et al.
Published: (2023)
WhAM: Towards A Translative Model of Sperm Whale Vocalization
by: Paradise, Orr, et al.
Published: (2025)
by: Paradise, Orr, et al.
Published: (2025)
Sound Field Translation and Mixed Source Model for Virtual Applications with Perceptual Validation
by: Birnie, Lachlan, et al.
Published: (2020)
by: Birnie, Lachlan, et al.
Published: (2020)
MT-HuBERT: Self-Supervised Mix-Training for Few-Shot Keyword Spotting in Mixed Speech
by: Yuan, Junming, et al.
Published: (2025)
by: Yuan, Junming, et al.
Published: (2025)
Few-Shot Keyword Spotting from Mixed Speech
by: Yuan, Junming, et al.
Published: (2024)
by: Yuan, Junming, et al.
Published: (2024)
High-Fidelity Neural Phonetic Posteriorgrams
by: Churchwell, Cameron, et al.
Published: (2024)
by: Churchwell, Cameron, et al.
Published: (2024)
MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio
by: Li, Qingcao, et al.
Published: (2026)
by: Li, Qingcao, et al.
Published: (2026)
Searching For Music Mixing Graphs: A Pruning Approach
by: Lee, Sungho, et al.
Published: (2024)
by: Lee, Sungho, et al.
Published: (2024)
Sound Check: Auditing Audio Datasets
by: Agnew, William, et al.
Published: (2024)
by: Agnew, William, et al.
Published: (2024)
Cross-domain Neural Pitch and Periodicity Estimation
by: Morrison, Max, et al.
Published: (2023)
by: Morrison, Max, et al.
Published: (2023)
Fine-Grained and Interpretable Neural Speech Editing
by: Morrison, Max, et al.
Published: (2024)
by: Morrison, Max, et al.
Published: (2024)
Maximum Likelihood Estimation of the Direction of Sound In A Reverberant Noisy Environment
by: Mansour, Mohamed F.
Published: (2024)
by: Mansour, Mohamed F.
Published: (2024)
Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?
by: Deng, Qixin, et al.
Published: (2025)
by: Deng, Qixin, et al.
Published: (2025)
Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
by: Kong, Zhifeng, et al.
Published: (2025)
by: Kong, Zhifeng, et al.
Published: (2025)
Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking Techniques for Speech
by: O'Reilly, Patrick, et al.
Published: (2025)
by: O'Reilly, Patrick, et al.
Published: (2025)
Similar Items
-
Generative Audio Extension and Morphing
by: Seetharaman, Prem, et al.
Published: (2026) -
Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations
by: García, Hugo Flores, et al.
Published: (2024) -
Audiocards: Structured Metadata Improves Audio Language Models For Sound Design
by: Sridhar, Sripathi, et al.
Published: (2026) -
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
by: Kumar, Sonal, et al.
Published: (2024) -
Video-Guided Foley Sound Generation with Multimodal Controls
by: Chen, Ziyang, et al.
Published: (2024)